Pith. sign in

Paper Citation Record · LEDGER

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

As of 5 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 100 inbound Pith citation observations for arXiv:2601.11868.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.11868 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T03:37:07.841385Z

measured 120 of 120 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 217 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T07:17:13.420563Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved8
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch4

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ba6d4cc9-f1fd-4bd6-8dcc-cb5f6381aad1 · outbound

This paper cites TextArena.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces TextArena

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:37:08.016661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:27083761712365659c007434f858d546d0ccec5ebf1598f6fc8d7d6c737cd3ee

Observation cd9c0420-6168-4263-b307-9dde5ac8dbb6 · outbound

This paper cites VisualWebArena: Evaluating multimodal agents on realistic visual web tasks.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces VisualWebArena: Evaluating multimodal agents on realistic visual web tasks

Reference 2

Resolution
metadata mismatch
doi, observed 2026-05-11T03:37:07.976590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:ca20d5a80b9b915794d4ea15ed5f3420b73a2a1d496d6b96ee9a65c3365cac1c

Observation b485e30e-fadf-47a8-bb07-efe7236bdf3b · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces gpt-oss-120b & gpt-oss-20b Model Card

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:37:08.005699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:c77c8b509aee7d75c9b211d835be6d41cc6d5f0ff5f874edc60d98a78978aacf

Observation 33ee05cc-1213-40ed-b33c-c6a79eff3c2c · outbound

This paper cites K., Krupke, D., Kidger, P., Sajed, T., Stellato, B., Park, J., et al.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces K., Krupke, D., Kidger, P., Sajed, T., Stellato, B., Park, J., et al

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:37:07.967629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:7dedaa7962e909bcbd9f7e093e139d80e0c4c987fa331fc17ce258497df8cf66

Observation 66ecafb5-6e96-4e31-a056-9461b2744708 · outbound

This paper cites Establishing Best Practices for Building Rigorous Agentic Benchmarks.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:37:07.992049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:03533c20dbd79eb633b52704382c63036bd5b3ef9966f1e5523f0523587d62e5

Observation 95239b1d-5d3b-44fe-8c52-63b01aefd6d2 · outbound

This paper cites 18 APPENDIXTABLE OFCONTENTS A Detailed Results 21 A.1 Comprehensive Results.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces 18 APPENDIXTABLE OFCONTENTS A Detailed Results 21 A.1 Comprehensive Results

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T03:37:08.114457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:9c7060e1f28344ad25e01bbaf2908136731d2539b5ae754b8c94fbec31af6c42

Observation b7a596a1-1c32-4bbe-8ce8-93c5f9ed1886 · outbound

This paper cites an unresolved cited work.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-11T03:37:08.134349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:c3d01d2b275b5d9bc31bb039ece6702b62b164b70f129ea89906f728135a7d11

Observation e3bb58bb-4a28-4859-8580-3b1f89c34848 · outbound

This paper cites an unresolved cited work.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-11T03:37:08.142535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:bd5a77c983e40b77284821a61fbfddce826f43355c5863fc2e7324e568fd812f

Observation 5f12f90b-f0f7-4a08-a6f0-9400bed77132 · outbound

This paper cites an unresolved cited work.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-11T03:37:08.146944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:732993c040d84c32101ca6386742376b3a011401774f5ffcfbe7dcf109ba34d2

Observation 1ea20868-973d-44d0-b83f-5744ccd0f5de · outbound

This paper cites outcome".

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces outcome"

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T03:37:08.028130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:c3bfde0e15424e2c5f8d23f9bdcdc9e11a4a3c7f85024f9945df73de865c0332

Observation 704a83f3-364a-48de-bd1a-fee2d68af6f1 · outbound

This paper cites an unresolved cited work.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-11T03:37:08.035125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:c5589aaef3c9c053798da6a7527f6222d61809d55aaa3be140e31ea16f725805

Observation 18234bb5-8c76-4752-91b3-3aa121ccce72 · outbound

This paper cites Independent inspection of the released tasks con- firmed that they are well-specified and largely free of ambiguity or underspecification.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Independent inspection of the released tasks con- firmed that they are well-specified and largely free of ambiguity or underspecification

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T03:37:08.042998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:29992c3604b6e442b4dd0eaf883a51266a15b753795bc65090d1d7aa0ea031fa

Observation 69736d52-131f-43c6-8a5e-72c1743023fe · outbound

This paper cites held-out.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces held-out

Reference 13

Resolution
malformed identifier
raw_fallback, observed 2026-05-11T03:37:08.049443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:8b33478ade0c08dfc919543841fa9ab3186c7e13fb4ccce92dd6a7001fd281f2

Observation 8420948f-997f-420e-8658-33e85060d913 · outbound

This paper cites an unresolved cited work.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-11T03:37:08.055164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:21cd491c0bef2c44f6cc8158d29753247edfb51673756f43da8c350c69a8176c

Observation 79ae3ee4-b721-4884-a081-76231d0e48a3 · outbound

This paper cites We’re looking for indications that running the command has failed - carefully analyse the outputs to determine this.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces We’re looking for indications that running the command has failed - carefully analyse the outputs to determine this

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T03:37:08.061385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:84c26b756b01889c5af99ab33dfb7df95755b97b168c9a846a64c7bc1c9bb14d

Observation b78106c3-ece6-4f01-8d47-5bb24889ec82 · outbound

This paper cites an unresolved cited work.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-11T03:37:08.068234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:4622ff455a587d742abac6e65ffe3918fa558b05a1db60334ac73ebcfe0640e2

Observation df70fb1f-1c80-445d-b7c3-83b63bab7be6 · outbound

This paper cites an unresolved cited work.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-11T03:37:08.074051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:2bdc6bd5a398f37e320a09d029770b9b9cfb8b7e7b9e6e07e8dedbf1f907ece4

Observation 60003a15-e549-4f65-bdbe-2021236c8a53 · outbound

This paper cites You are about to hand off your work to another agent. Please summarize what you’ve done so far.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces You are about to hand off your work to another agent. Please summarize what you’ve done so far

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T03:37:08.088349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:6a61ce3e0575c9db674770fe710a54b9a732bcecd309bc32d6245372a15c30d9

Observation c92b3ab6-065e-40ca-9339-c9cd25485b5a · outbound

This paper cites an unresolved cited work.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-11T03:37:08.101181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:c003cde72febab40942f662c36bceb34029896064ee74b9cfd0a637c9c4693ae

Observation 75a8499d-5053-419f-8936-085fa367be89 · outbound

This paper cites Results: X Y Z.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Results: X Y Z

Reference 20

Resolution
malformed identifier
raw_fallback, observed 2026-05-11T03:37:08.109352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:d5d28d1e6af585ae4d533ae8e0d1776e10dff7152eeab7ba8e11eda3e4337016

Pith citing papers

Observation bc2c7d59-08bb-43f3-8808-9fbaa0217091 · inbound

SERA: Soft-Verified Efficient Repository Agents cites this paper.

SERA: Soft-Verified Efficient Repository Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T07:17:13.420563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:17:13.420563Z digest=sha256:42c395c0980f08b57a0cb25a2411d72995d87b1646bde7ee71c6c15dd2c10f8e

Observation 3cadcd68-2d55-4d6e-ac8d-6a4aa8a99808 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:7c9e5d860bff6805ad1faf562f216981fd20cf9c43090b83488b5c526b6f2d1e

Observation 8d070eb9-7a96-4dbd-8588-290400dabc58 · inbound

Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation cites this paper.

Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T03:07:11.985780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T03:04:17.755968Z digest=sha256:bb818ec6782f31134c346764ec62634e7e4afad6a95cbfb4ef514a061cfc13b4

Observation c2502f82-f2f0-4f31-bce7-5ad70f6b8715 · inbound

VeRO: A Harness for Agents to Optimize Agents cites this paper.

VeRO: A Harness for Agents to Optimize Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:06:30.926621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:01:50.547095Z digest=sha256:5859be75b3b244be4fe7d0194cce72f824a52b9898918121fbee9da0b068fcce

Observation 3f66d6e5-9cbe-41ef-9b2e-302297c6924d · inbound

VeRO: A Harness for Agents to Optimize Agents cites this paper.

VeRO: A Harness for Agents to Optimize Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T20:45:12.321842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:45:12.321842Z digest=sha256:2aa941885cbd926cb6a8cb0bbe6169f1b8269f26daf36d1e1fe55d15273954f7

Observation e43ae4b7-e82f-472f-af09-1d853d86e373 · inbound

Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development cites this paper.

Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T16:00:10.128726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T15:59:29.910200Z digest=sha256:fc2a4656a494a335b3058fb0df55e3db58998a6da2e379befc58526ebcd96010

Observation c361e55c-5065-4bb8-bfcb-60f76da75665 · inbound

Effective Strategies for Asynchronous Software Engineering Agents cites this paper.

Effective Strategies for Asynchronous Software Engineering Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T20:49:09.477849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:49:09.477849Z digest=sha256:b52924028be76511c6f8773d36a19cedea6b3ddd89aefcbbafe9f865b6cad028

Observation 50b1f0e6-6b95-4ac1-a3fe-226c9e4e69ea · inbound

SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks cites this paper.

SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:18:21.683940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T00:15:40.442807Z digest=sha256:75c5a138fc47c81b1d3b3421185e6bdaf7d0fbf14d8293f1292738c245692961

Observation c9e2e6ae-129b-440b-8428-c68864acdeec · inbound

Meta-Harness: End-to-End Optimization of Model Harnesses cites this paper.

Meta-Harness: End-to-End Optimization of Model Harnesses Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-13T16:15:58.227250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T16:15:57.877354Z digest=sha256:a14ed45d37e4d0ce6a08108840babf57d1c9e5ad6c88301f11c2b5200ce175a9

Observation 60c625e6-06e4-4417-aa07-988588164588 · inbound

Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills cites this paper.

Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T09:29:51.204678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:29:51.204678Z digest=sha256:20c1458024f5735b2a1a0b6ed8eecabf055e55f4fc85978cf0ab94420e648bb8

Observation b16f0811-0ba8-4acc-89f3-22f44b099f5d · inbound

TRACE: Capability-Targeted Agentic Training cites this paper.

TRACE: Capability-Targeted Agentic Training Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T20:12:06.752461Z digest=sha256:e816a52da5790b0c7d66cb2d792381596b570f3ad70cb5c2180719d692f73ee3

Observation 2c22284f-4079-4a85-aa42-a9581bc39753 · inbound

AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments cites this paper.

AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:07:46.077831Z digest=sha256:953d3c3c5072362c524c73ce50e08dae3f03e67edf8c208336b0723851c95b51

Observation a383907e-af74-40bd-9195-d767b431b951 · inbound

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents cites this paper.

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:59:59.383781Z digest=sha256:6d5821f1f4e76acf76c48a4a5c7fc9794f8d4ad0c466d41e89255f53d0597e5b

Observation 9d40e8b6-cfd2-4416-a969-90a908bbc508 · inbound

Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios cites this paper.

Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:35:03.746555Z digest=sha256:6810f9b857956b8f1caa8bc992f044946bb98c6f9d8f1830a10b1322d2922d84

Observation 4818b396-6ccc-4d60-9eec-fce92bdba149 · inbound

Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios cites this paper.

Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T16:42:14.837132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:42:14.837132Z digest=sha256:c935362adcd4022a40772c1eb6155ba70377a9e2af1537f5b84add203140a37e

Observation ecd9eaa4-0a65-46c0-a623-229feafe242a · inbound

How Much Heavy Lifting Can an Agent Harness Do?: Measuring the LLM's Residual Role in a Planning Agent cites this paper.

How Much Heavy Lifting Can an Agent Harness Do?: Measuring the LLM's Residual Role in a Planning Agent Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:18.728493Z digest=sha256:ba5d10a91808265b2b2f6e8eb37288b885061e10cc6d9bf7f7d6afcbed47efb5

Observation 01afe455-1add-420e-9124-63c932f63770 · inbound

Prediction Arena: Benchmarking AI Models on Real-World Prediction Markets cites this paper.

Prediction Arena: Benchmarking AI Models on Real-World Prediction Markets Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:58:14.109358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T22:56:58.482896Z digest=sha256:79d4825493b5183c5d7aad35d5e6299257889617147ce2b928457df30e53b53d

Observation f81e0b6f-b431-4ec2-808e-eeb124a87ee5 · inbound

COMPOSITE-Stem cites this paper.

COMPOSITE-Stem Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:41:52.868766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T17:29:15.163053Z digest=sha256:8ca2fe28ae2b826f3cb1e79b29b1c64e113d32803140ee1aba4288caf7aaaa55

Observation d1cb947c-175e-4db4-9988-9380e13e9586 · inbound

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation cites this paper.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:57.808333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:e313c4cea7a05a5016a4128e3a8d738e10d8b7e6ae7e2c1a6779c2c704c983fb

Observation 31d7448c-c4a4-4761-b50f-089bd20710e4 · inbound

From Context to Rules: Toward Unified Detection Rule Generation cites this paper.

From Context to Rules: Toward Unified Detection Rule Generation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:55:59.987338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:25:59.609473Z digest=sha256:b49861cb093e64eac0c141d94239977810a925191e478df190de35dfec07a652

Observation 24ac4b31-03b1-4834-8f89-731d838722dc · inbound

From Translation to Superset: Benchmark-Driven Evolution of a Production AI Agent from Rust to Python cites this paper.

From Translation to Superset: Benchmark-Driven Evolution of a Production AI Agent from Rust to Python Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:46:08.268286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:50:05.278161Z digest=sha256:8dac3f0344eaf167aaa541eda52e994a4f0503fe9dfa78f7d5e87b2a86a7fe67

Observation b4eb09ff-74d5-4d1f-a6e4-066780d33ace · inbound

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization cites this paper.

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T09:05:57.732149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:17:32.290531Z digest=sha256:91e20fce78b9d8e37962d16ed70ccf0da5ae8416f35a575f48c76d48a69f321c

Observation 9b070afe-5685-4a0c-87fd-f39ebfa076c0 · inbound

Towards Long-horizon Agentic Multimodal Search cites this paper.

Towards Long-horizon Agentic Multimodal Search Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:06:00.175925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:40:32.137708Z digest=sha256:fddae377aaedc715bfc4fd1b09a87fbf5da3ed80f2834e7a988b9699207422cc

Observation 342b5f46-0949-417f-b797-6b2d68b7ad42 · inbound

Exploration and Exploitation Errors Are Measurable for Language Model Agents cites this paper.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:21:02.382824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:ce1db230f7bffade7ec0018581ee6aac4a4fb60385f1854cc5c5892052ceabca

Observation ce1cea9c-d44b-491e-86d8-291b39ac4e5b · inbound

Don't Let AI Agents YOLO Your Files: Shifting Information and Control to Filesystems for Agent Safety and Autonomy cites this paper.

Don't Let AI Agents YOLO Your Files: Shifting Information and Control to Filesystems for Agent Safety and Autonomy Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T12:25:50.595161Z digest=sha256:d322b1c918225112aa977757d7b8a39dd7707852904976c75a7a0d27c2ee944c

Observation d6db8f4c-b25d-40c8-9fc5-d247cdfbb397 · inbound

HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks cites this paper.

HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:27:11.257817Z digest=sha256:a8d53bc631a0a5b6fa5de62dad5151af930b40d72cf1864a62bce98c08a79e6d

Observation e22b2236-bded-4744-b39e-beffc944a32f · inbound

KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving cites this paper.

KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:58:36.525442Z digest=sha256:c3ec175d348477db037e52b3a0aee7b803fcf36ab70e6ca498124dc27a5c2d1d

Observation ebf058fe-d98b-4a30-a5a3-ba495e4e3a9a · inbound

BranchBench: Aligning Database Branching with Agentic Demands cites this paper.

BranchBench: Aligning Database Branching with Agentic Demands Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:12:46.899179Z digest=sha256:11cb22e76e08a28724a4c77ef0a886357319159e7284f560a0ac76b0b20f95b5

Observation 222dc975-e5bb-4afa-974d-e8eea22f129d · inbound

SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents cites this paper.

SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:13:32.434201Z digest=sha256:ad30bd54afa30975bc5aaff982cd19a646cd65f85b28c4781895b192bc0358f4

Observation b3620eee-499e-42a7-a8df-cab7d7b76c02 · inbound

Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories cites this paper.

Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T05:48:44.687520Z digest=sha256:873fd7f8cac7534f7f305823a6a549ccf459f805cbe0e24189444dfb6309f1ab

Observation 73a253dd-f2fd-451e-b667-631dd09377f5 · inbound

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence cites this paper.

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T05:24:00.503836Z digest=sha256:dbede02c09187f0dee47b01d2fbb8429d1a9bf459307d4cea9fd00866fc9d13f

Observation a4525fe0-769f-415d-9c41-d2dbc0339314 · inbound

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents cites this paper.

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:10.854057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T06:36:52.202847Z digest=sha256:41fb059956715b690fa787d9855390fd1a56abeb27a93c92d24339dd11399f64

Observation 8c5651dc-ed4c-4bbb-8999-66f412d579d6 · inbound

Toward Scalable Terminal Task Synthesis via Skill Graphs cites this paper.

Toward Scalable Terminal Task Synthesis via Skill Graphs Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:51:16.809578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:13:50.484665Z digest=sha256:ebeded8e67950fb1b497454b06ca29aa17ef730e996d8dbe9333d2e5b9dcdd0e

Observation 96afacb2-6ed8-4f57-a9e4-1b1a3d06d911 · inbound

Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses cites this paper.

Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:46:45.190352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:18:46.514331Z digest=sha256:d748a80a7d6d8296c47abf85047506840dccd74876c2ac859b5c8ea55d097b93

Observation a92fa150-3bf8-4f25-b86a-c31343bdc321 · inbound

Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses cites this paper.

Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T23:53:51.714728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T23:50:56.311572Z digest=sha256:e7c48744c4496bc689c79ddf67e73cb180299ccd1c4afc7f3c9a9e116f64950e

Observation 0fac8e12-f199-403e-8617-2cb5f533c068 · inbound

MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems cites this paper.

MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:30:56.899306Z digest=sha256:6317bf4201e14ad6f5a24989fab92584897c6c629071830d0b31d6ce8b9da546

Observation 2f684907-36ef-4aff-a4df-fb2158e37a91 · inbound

Heterogeneous Scientific Foundation Model Collaboration cites this paper.

Heterogeneous Scientific Foundation Model Collaboration Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 107

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T08:50:05.980191Z digest=sha256:3d89f1135a416cc7eab9ffa6fdf69e01a0995146776463220b0dbdabfd9c9298

Observation 1b7cf7a6-cd43-4d3a-949f-0897f251da6e · inbound

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design cites this paper.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:36:30.975678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:c4c8d86e6f2c71b2a7987b27955c71a675c0781831b37e8aa6e13e646624cc7a

Observation 9546a39b-4954-431b-9a4a-4abc9fc99d66 · inbound

Crab: A Semantics-Aware Checkpoint/Restore Runtime for Agent Sandboxes cites this paper.

Crab: A Semantics-Aware Checkpoint/Restore Runtime for Agent Sandboxes Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T10:31:28.618912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T05:56:11.552190Z digest=sha256:d7d8b3951539c396fe7337ab12d28ffb1a960142ad1fc7656933735c7410cfa5

Observation 6bdf56a8-1aed-4386-ae4d-fb2556d38d18 · inbound

LLM-Oriented Information Retrieval: A Denoising-First Perspective cites this paper.

LLM-Oriented Information Retrieval: A Denoising-First Perspective Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 126

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:01:19.839855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T18:54:06.144968Z digest=sha256:e43c4157a14aac352a3c952c837e381b8f4297dd57bfa9e3f3973f816e48bebb

Observation 1c02537c-702a-425d-9fc1-fb4fe28c57a1 · inbound

LLM-Oriented Information Retrieval: A Denoising-First Perspective cites this paper.

LLM-Oriented Information Retrieval: A Denoising-First Perspective Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 129

Resolution
verified exact
local_arxiv, observed 2026-05-21T00:19:16.648530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T00:18:32.423103Z digest=sha256:474eceded70db0b7aadfe9b9eaa11d4335fe822b9faef458e3bebcba3a2593cd

Observation 961c9605-9309-40bd-9df3-2f434e68e2d8 · inbound

BUILD-AND-FIND: An Effort-Aware Protocol for Evaluating Agent-Managed Codebases cites this paper.

BUILD-AND-FIND: An Effort-Aware Protocol for Evaluating Agent-Managed Codebases Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:26:12.626586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T08:57:41.818402Z digest=sha256:3c2994875f25c79e682bd6a56307300b33b67870af70300635586785a5da8ff1

Observation 6c3523ec-d2b4-49d8-baf1-3470940b2b51 · inbound

PrefixGuard: From LLM-Agent Traces to Online Failure-Warning Monitors cites this paper.

PrefixGuard: From LLM-Agent Traces to Online Failure-Warning Monitors Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:16:10.832441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T09:44:55.593277Z digest=sha256:3b7f1ec214d21cd276672d15e495cdf0f4b7d95fa9521bdba92607bd730f4dfe

Observation 0ebd8258-2635-42b0-acfa-36a76a63ddb0 · inbound

TeamBench: Evaluating Agent Coordination under Enforced Role Separation cites this paper.

TeamBench: Evaluating Agent Coordination under Enforced Role Separation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:00:55.485805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:55:51.358828Z digest=sha256:dc54893e916452c440bf80598d4c79e1bcfca021289c03be803d51f2893c49c7

Observation bb9829ac-ef4d-47a2-961c-d4ea2a3aded4 · inbound

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios cites this paper.

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:20:32.528550Z digest=sha256:b47298fe2682b0aa16c80834dc9570075b9440494781b10266fce126e40d1110

Observation 61cd7e8a-a615-4554-b1d0-b48146f8d843 · inbound

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios cites this paper.

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:49:29.207918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:49:16.239350Z digest=sha256:134802b9d8cb2c094a9389ca113b938d88f2dd76e6c72ad7fd1897715f51738f

Observation decc0cb2-5d3a-4d38-bb4c-c0ce013a4088 · inbound

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios cites this paper.

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T00:18:43.017866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:18:43.017866Z digest=sha256:07e54d8a42f4e7077c6273f4d8137931e3fab19f81e2a1bef78ebe1b7597cc26

Observation 9aebe349-74e9-4a5e-8932-3b0870b4ed46 · inbound

Learning CLI Agents with Structured Action Credit under Selective Observation cites this paper.

Learning CLI Agents with Structured Action Credit under Selective Observation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:37:08.148477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:59:26.100818Z digest=sha256:f9e99cbaf2dabfdb1726ef2aa9cc1c724fd94225780d246d4a5ece2bd090d5ed

Observation 9460b06a-2718-4a8d-824e-c7412ab0b561 · inbound

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution cites this paper.

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:36:58.326900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:28:07.557119Z digest=sha256:bfed4146ea6db0843a8c72f663df2093a3183d509cb80e3d5f4330e116d98fc7

Observation dcafdf38-1693-42ab-86cf-5468dfba4fc0 · inbound

MDGYM: Benchmarking AI Agents on Molecular Simulations cites this paper.

MDGYM: Benchmarking AI Agents on Molecular Simulations Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T03:16:18.565184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:2606f4bd16ed1235b06af2d1a9e682d33ccc1d1ad927591ee52cde89d42e260c

Observation df52b1ab-c41e-4b5a-a125-ec598d19c1a2 · inbound

LLM Agents Already Know When to Call Tools -- Even Without Reasoning cites this paper.

LLM Agents Already Know When to Call Tools -- Even Without Reasoning Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T02:46:19.060948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T02:44:00.329276Z digest=sha256:b3dc5bbf09787f8ee2b440b00be18cff1342a03401174f661c44b2062408cb1b

Observation f2f9c64f-d445-40d5-a80f-22e192a3ff2b · inbound

LLM Agents Already Know When to Call Tools -- Even Without Reasoning cites this paper.

LLM Agents Already Know When to Call Tools -- Even Without Reasoning Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:56:25.965009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-22T10:55:14.037226Z digest=sha256:16dd445414fac9924644892b8eefeb5f66dec24318884c76ed9e1fe761843293

Observation f3f32258-9907-456a-953e-b8350f686e5c · inbound

Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values cites this paper.

Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 83

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T04:31:21.357721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:22:06.172150Z digest=sha256:5efd952d4b0e9c7ee45bcadbad0400791595300a80cf7ace357c3c7ea6d57637

Observation e730d03d-012e-46db-ba61-4682e42e3a8a · inbound

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability cites this paper.

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:41:21.856981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:41:15.286881Z digest=sha256:fc10a57a7895fb53780d8824521ecdfa1fed0f6f64f545521dd930478d282293

Observation bf81f40f-85c9-4ef1-b7a0-a9afd75ccfb4 · inbound

WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation cites this paper.

WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:23.964559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:40:00.725327Z digest=sha256:dff4b6c380560963e5732f4df2f5df87b915ba8ed1dbbad8226523d1e640311e

Observation 6ff4237c-23a0-432b-b0dc-b69a0e5b4df1 · inbound

Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces cites this paper.

Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T07:16:31.018863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:29:33.497561Z digest=sha256:e1ee6acab2d9486100c00503d715f683ae23c8fbae54cde2a283cfd48ce5218e

Observation fca59e70-5b37-42a0-81e0-55231c125552 · inbound

Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces cites this paper.

Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T22:25:06.838406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T22:24:26.528532Z digest=sha256:a5b9248c8ae56fe353c56d0215e76d8e60a30fbc52933cae7ac1e79449c1a82d

Observation 51e3bf8a-4d6a-40bd-827e-3b03fabf1c24 · inbound

MMTB: Evaluating Terminal Agents on Multimedia-File Tasks cites this paper.

MMTB: Evaluating Terminal Agents on Multimedia-File Tasks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:07:00.481516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:03:22.390574Z digest=sha256:3c16d4dfdc7c80de4a284c6d98fb51715b1cd8a6d66a5c6c31f52c06748d5403

Observation a5c385cf-7725-43df-88d9-56073daa5181 · inbound

gwBenchmarks: Stress-Testing LLM Agents on High-Precision Gravitational Wave Astronomy cites this paper.

gwBenchmarks: Stress-Testing LLM Agents on High-Precision Gravitational Wave Astronomy Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T02:17:06.452705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:14:48.091245Z digest=sha256:836ea8d8c91b566849fc28707cce27697f4181f29ee310e4ae93960cf10b2156

Observation c32812f1-7fb8-42ee-940e-4ce580779882 · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:32:56.949393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:7ce1fc543c94810d177914f5a0923ad4a1e64116ab68895b72d4bd7bb62d523d

Observation 503b79ae-787f-49fa-83da-4f68cb2e5327 · inbound

AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation cites this paper.

AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T18:52:35.464751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T18:51:06.379266Z digest=sha256:3b1068456dc5ff232e2d76cfe433791d3a75b1af3686d5885f903e5e11d47e8e

Observation 15c81c16-6fee-4891-adf5-99f8d6bc80d6 · inbound

PREPING: Building Agent Memory without Tasks cites this paper.

PREPING: Building Agent Memory without Tasks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:15:06.120258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T06:14:14.385586Z digest=sha256:ded3cd4162c1901c6f373103e992f66b4966158b4428cdf09eec7c7f9e5c7f41

Observation 309e8e6a-3dae-44f3-8ced-fbfd0fda3b69 · inbound

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction cites this paper.

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:05:06.740995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T06:04:03.605898Z digest=sha256:624f72fb08210b452b364bd129ac23183a9f1d9f1af7d9279c8742ff50a7bf27

Observation 4fc34411-3e50-43d2-ba18-160b02c5230f · inbound

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing cites this paper.

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T04:49:44.564961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T04:45:39.641535Z digest=sha256:9659140d1c44b231524d8585660d46c0af8f475335ea8991bf7c6835893388a0

Observation aed8048e-41c0-4af3-892a-79f4cf2aa89d · inbound

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing cites this paper.

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:15:04.197628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T21:12:37.353935Z digest=sha256:0a3263ea90bdb0ece4138f4e4f527b7e9ab103d2359d35b32876a175915880fc

Observation 78cfac98-5e6c-4228-a650-34cfaf220276 · inbound

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents cites this paper.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T05:09:45.015967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T05:08:15.750558Z digest=sha256:945b386d25a65e1438b9d2a050853cf03e87e1e9fbbb239dad0b7cd87d1144cb

Observation 6ea8b4e0-dd3e-471f-9a61-85c3ec0be6e3 · inbound

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents cites this paper.

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:23:43.298529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T20:19:21.824216Z digest=sha256:5b0c7300dd513dbe57caff05fb2d71ced1cdb82cdd1c51a6f9aec0c380a62036

Observation e219b491-3529-4255-91c5-73f8bd914e35 · inbound

Do Coding Agents Understand Least-Privilege Authorization? cites this paper.

Do Coding Agents Understand Least-Privilege Authorization? Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:37:40.123294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T16:34:14.379419Z digest=sha256:3b588091a80ad6b6138ce6db0e4270c90b2fc6a889f91f7eb2c2dbc70a6c782c

Observation 2f93dc2e-a6e9-43c4-845e-e4d0be8ab33a · inbound

Orchard: An Open-Source Agentic Modeling Framework cites this paper.

Orchard: An Open-Source Agentic Modeling Framework Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T09:51:21.271005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:94d4b3e42d6fb771311e740127fd84dc76e77443664002173e8091ff67925a97

Observation 88082ec2-1c15-4900-a813-9175ecfadaf6 · inbound

Orchard: An Open-Source Agentic Modeling Framework cites this paper.

Orchard: An Open-Source Agentic Modeling Framework Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T14:05:15.131258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:05:15.131258Z digest=sha256:5d0bc5d4c1bc3a4e1854d01dba3d7bc3b6cc1cc264f2ef9f4d6a3b7e96f81da2

Observation 27abc1ab-56ef-4fc6-9416-469e0595125f · inbound

The Scaling Laws of Skills in LLM Agent Systems cites this paper.

The Scaling Laws of Skills in LLM Agent Systems Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:13:37.642484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T18:10:08.737710Z digest=sha256:a4ed89e9b5380ecdaad2e3edf2caa40009dd458ac73e66c9f3b663ed9dcd6f20

Observation a69b9c14-c1b8-4e96-acff-34dcef2fff66 · inbound

R2V Agent: Teaching SLMs When to Ask for Help cites this paper.

R2V Agent: Teaching SLMs When to Ask for Help Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:28:59.778484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T20:26:29.804427Z digest=sha256:958b4f72e24244ae18513f2f1417b159f50d5ac4897a548fac70f4fa11786384

Observation 75a26867-59a5-44f7-9907-2686bbfb54e2 · inbound

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? cites this paper.

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-20T17:48:48.854593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T17:45:02.896703Z digest=sha256:6c37667a26be8cfeef9f309c0c904f1cd01154a426894e0f1445400abfabf5ac

Observation 715c29b0-5284-465f-88ec-2763e48ffdd8 · inbound

Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench cites this paper.

Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:33:25.519246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T15:31:25.079191Z digest=sha256:a4eef380eb12303629016fc98e1e8a2fd373a33e66625915b83bbe763b1c727f

Observation 45f591ad-0cee-4f48-871a-0c59c6c30f26 · inbound

Responsible Agentic AI Requires Explicit Provenance cites this paper.

Responsible Agentic AI Requires Explicit Provenance Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:13:21.291811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T14:09:17.469239Z digest=sha256:589db912c17406d425a13ce21d93bdf07dcc4d2c9af96e4a10a9c719925e3d46

Observation 532443f7-28e3-4935-ac72-dc09c9cf633a · inbound

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games cites this paper.

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T12:28:17.276478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T12:24:06.062957Z digest=sha256:290106d86ba5f57bed635a9762b2c772714078d6aad9133f871e11d22065dc72

Observation 93d737d9-6e7b-42b8-aaeb-c649ceb56585 · inbound

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games cites this paper.

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:45:23.215403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T05:45:04.573722Z digest=sha256:b172b785f590a468cc1b44f1fcb3eb93a3079e7d7a6524c4f6e924ccdc48e370

Observation ce1150a4-67d4-4813-9d06-6f50130693db · inbound

SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution cites this paper.

SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:43:15.124799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T11:40:45.397038Z digest=sha256:e6c4b45c97a6e553e32e076ea87089f74769badd2c2abe2c5ad8e86be29eb067

Observation e97fe467-d648-43b5-b617-064f30d817af · inbound

HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models cites this paper.

HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T05:53:04.302217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T05:52:53.880521Z digest=sha256:a4ba17a689da853260edcd9bf07b8cc341d837314bc0f37a670c28f916c32686

Observation 3cc47777-f3ce-4c2a-acd6-9db4470922e5 · inbound

Open-World Evaluations for Measuring Frontier AI Capabilities cites this paper.

Open-World Evaluations for Measuring Frontier AI Capabilities Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T06:39:43.684865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T06:38:51.427985Z digest=sha256:d8fa9ff3ceae6c8d283c12a1734541f3272eb33ad77ff3afe65e6a260cea6810

Observation c3805acd-f5f9-4626-8013-2009cb555e79 · inbound

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale cites this paper.

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:59:45.555982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T06:56:27.532299Z digest=sha256:9b434b3a9d7f61267cc6d668855024608827f6ae645ed68c34b1c0a84a26409a

Observation 9603e92f-7fe6-420b-985b-9c7047c38e74 · inbound

Terminal-World: Scaling Terminal-Agent Environments via Agent Skills cites this paper.

Terminal-World: Scaling Terminal-Agent Environments via Agent Skills Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:54:35.887170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T04:54:17.662339Z digest=sha256:3ab62fc890802da148ef2f40820c98676867abf0f3cb47bb3029a8e21e69698c

Observation 0548592f-515c-49c2-a65e-efb64215ce0e · inbound

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents cites this paper.

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:23:57.572966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T04:20:43.780849Z digest=sha256:3ab52b2ad8fcb3e24b23cf3850912a2e52fbc157c6a9147cceeb260fb990e14f

Observation 4c573288-681b-4aeb-a22e-3feb1f2a2278 · inbound

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents cites this paper.

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:51:21.850352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T09:46:26.124683Z digest=sha256:c618480492287bf0f0c0206395ffd76b0961f7d7332c4c5bf5e56860bc027fe3

Observation 68cb9568-1160-4b29-bea1-9285dd28addd · inbound

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents cites this paper.

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.670499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T17:31:25.538977Z digest=sha256:f07d5eac0f70cd346eb1522ab597bac02e988365b94310c4073045abfdbda841

Observation 5ee85988-18d5-4098-a1f8-a56ac7e999f1 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:51:07.764040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:3a4fd2314c7f53c2b7f5a101a8ab4018b61c8dbbca4a11fba90088c154360561

Observation de70236e-354d-44c0-a4b1-0b1eeace607b · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:06:42.930484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:4ea777ceed201c6dfda168910e9dae2b1999766bf1ce80a079868ed0fc3b6690

Observation 12e2610c-841b-4ae3-a2c5-0ffe5acb0b16 · inbound

Stop Comparing LLM Agents Without Disclosing the Harness cites this paper.

Stop Comparing LLM Agents Without Disclosing the Harness Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:25:45.853016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T23:15:37.160073Z digest=sha256:b210883213f29528914269d0e495e17fdd4833ba121b3865ff0270b4830e68c5

Observation c6e275ad-ec3e-4325-90da-d6c1b7f0fded · inbound

EvoCode-Bench: Evaluating Coding Agents in Multi-Turn Iterative Interactions cites this paper.

EvoCode-Bench: Evaluating Coding Agents in Multi-Turn Iterative Interactions Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T16:24:53.896074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:7e2455a1ff6ae45458a6e14f5a41774eb39b0f6a05501a9c46c3272823daeec3

Observation 841dbef7-3a3c-43cb-ab9a-74663fcee331 · inbound

SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills cites this paper.

SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:24:55.867127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T16:18:59.083880Z digest=sha256:bc4576b53a177287f9ee9a94ef8edc54e2127fee05f407d54aec8a8943f91100

Observation deefa0d1-ddb4-4fb0-8e1c-bc03ee987e64 · inbound

DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations cites this paper.

DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:14:40.626642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T13:12:46.927103Z digest=sha256:3651898358c9ed7b2f2372a8c2a4bfc6b7ef2cc40d04a878684a520e5090c9d1

Observation 28ff39d4-3f31-448f-8496-ced88e58a6ae · inbound

Heimdall: Formally Verified Automated Migration of Legacy eBPF Programs to Rust cites this paper.

Heimdall: Formally Verified Automated Migration of Legacy eBPF Programs to Rust Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T22:04:00.140390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:02:16.597531Z digest=sha256:59c9faa6a2c594ea6d8378b17a8152cb8d0ef54bef43038b67cbac2544aef5a8

Observation 2a338a80-4f3c-4d43-95ad-779cfd6205db · inbound

From Model Scaling to System Scaling: Scaling the Harness in Agentic AI cites this paper.

From Model Scaling to System Scaling: Scaling the Harness in Agentic AI Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-29T21:33:59.178143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T21:29:57.326718Z digest=sha256:318a775db8840a81b2b21049c690c12b7a9e8932fc4d7258431f1f9ba3815cdf

Observation fe3a6ea6-693b-4610-a129-dfce44b43809 · inbound

Agentic AI Workload Characteristics cites this paper.

Agentic AI Workload Characteristics Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:13:58.843019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T20:11:50.722787Z digest=sha256:9f266b43e94a620cab0a829a27f7a21bae6e01a7dbf9e739fc893ba4b3049f2a

Observation e274befe-b446-46bb-b4a5-9234ca378836 · inbound

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems cites this paper.

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T15:33:32.776969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T15:32:21.737028Z digest=sha256:66e0d2a3fbfa9172dc9d0fb4ad45cb040e77010686d3e510d2e07cc9f1e70459

Observation bbb67e4c-f862-46d1-a376-9406eedf60ca · inbound

LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents cites this paper.

LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:33:13.901903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:27:45.349782Z digest=sha256:5c0d6b4f82e69aa7ecabf967a1d3b5cd220618b77c834b8df1a64fd9e13ca827

Observation 54a1d8d7-304f-474f-b75a-e6fb64032235 · inbound

TraceGraph: Shared Decision Landscapes for Diagnosing and Improving Agent Trajectories cites this paper.

TraceGraph: Shared Decision Landscapes for Diagnosing and Improving Agent Trajectories Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T19:36:08.589449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T22:22:54.473952Z digest=sha256:6241c729f01470036e30868bcb227cf417f9f1e397ec5af0e36232e4a061333c

Observation a0b672bc-4ad4-429b-8d7d-b5e3125865f7 · inbound

Stateful Online Monitoring Catches Distributed Agent Attacks cites this paper.

Stateful Online Monitoring Catches Distributed Agent Attacks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T19:56:11.080961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T21:54:44.072929Z digest=sha256:86a6cf48593b79f9306390243abfd23fd52f5e2045b9d0508550dbd87e45bab4

Observation c1ef4e1d-6e4b-4cb0-b36a-03adb6528df3 · inbound

ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use cites this paper.

ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:16:00.188334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T23:03:36.851403Z digest=sha256:9fd6b48765ba30f5a4d809cb6140f65cd21e5b830324285573edd340d3e66674

Observation 1404f36a-9048-437a-aee0-5f5b65eb72eb · inbound

Sakura: An Approach for Generating Complex Tests from Natural Language Test Descriptions cites this paper.

Sakura: An Approach for Generating Complex Tests from Natural Language Test Descriptions Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-06-28T18:42:29.432773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T18:40:49.564295Z digest=sha256:399dc83bb7c549f60cd76f4412fb775dfb7d54c1827f4b2f93482dae1f8ac95e