Pith. sign in

Paper Citation Record · LEDGER

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design

As of 4 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 1 inbound Pith citation observation for arXiv:2604.28093.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.28093 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-07T05:10:58.065201Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:25:08.066021Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact6
  • verified fuzzy4
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 433f73ab-51e8-409d-8b71-d0aea40d42b2 · outbound

This paper cites Concrete Problems in AI Safety.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Concrete Problems in AI Safety

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:36:30.967971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:49a8e19bb5b4924626b009e3a7d49c29590d7ea2ec151e550860ce17a577ec8a

Observation b94a18b1-ec0b-404f-afae-70e1f4292d29 · outbound

This paper cites Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:36:30.961371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:e8060f4f785610e0b0524671ddcc1ce4c04f3d19a88a17d9843681c1654a1a13

Observation 7bd7c221-8fb6-484a-97c1-c3ecd34cc25f · outbound

This paper cites an unresolved cited work.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-27T10:49:02.977699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:807466d976c405da44878c224db457e1d2d8234a0488dbbcc33b34eab534c06a

Observation 4d46291f-fd14-46bd-8ccb-5e1e21ef324d · outbound

This paper cites Dell’Acqua, E.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Dell’Acqua, E

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T10:49:02.983915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:20efd83267d0c2f8db01de182f643b1168f108ec9c4c5ebb35611c3d12138665

Observation b4af3104-1a7f-4deb-9d2f-1669321c1666 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:48acedc67db4e9888ba8490d49c44089c8c5f93356bf7589c22536f77cf81940

Observation f422a5c1-db34-40e8-8818-136a45f0867b · outbound

This paper cites Krakovna, J.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Krakovna, J

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T10:49:02.974312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:127021daeab45d881a2a1666bf664c487933098c4822faf8a09b2189c71f9b3d

Observation 3db5aac2-0534-4751-a061-16d7699ed47a · outbound

This paper cites Natural Emergent Misalignment from Reward Hacking in Production RL.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Natural Emergent Misalignment from Reward Hacking in Production RL

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:30.971759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:8487f043d3198180cb35d40451dc3082a933c75cc4fa4fb6496ac124d749c764

Observation 1b7cf7a6-cd43-4d3a-949f-0897f251da6e · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:36:30.975678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:93e08043423e78b9eea5c243a1585ecf820a1d02d3b558adbbc33eb627f0aadf

Observation 8a086f59-ba23-44d2-b88a-926bbcb348a0 · outbound

This paper cites Launching the OpenThoughts- Agent project.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Launching the OpenThoughts- Agent project

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T10:49:02.987471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:11a4b9710899f167d80b179c3a2b2dc875162e5f27e6e9a2ae67b49625f9d9d8

Observation 7e9de2b1-74b7-4445-abf0-287fde749d88 · outbound

This paper cites an unresolved cited work.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-27T10:49:02.967565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:1316e231fb9344a5cf15e6019543142d62e5d28ac605c16cae1815adadb7e4dc

Observation eb01340d-9a2c-4242-9c66-0f175b648b52 · outbound

This paper cites an unresolved cited work.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-27T10:49:02.990589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:b582531efc028e46ea0aec8cfcac4365d2d0b41c55125486c099435a4962ccef

Observation 31434c35-f1cf-4fd8-a5dd-cc7b48aa5e9d · outbound

This paper cites Von Arx, L.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Von Arx, L

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T10:49:02.980718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:5dd31f7ec834428e5aeeba201328017e6975fdabf008a9eed24d0776189fbe6d

Observation e9ceb1b8-34f6-46a6-8db0-6a882ed17d2a · outbound

This paper cites Let it flow: Agentic crafting on rock and roll, building the rome model within an open agentic learning ecosystem.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Let it flow: Agentic crafting on rock and roll, building the rome model within an open agentic learning ecosystem

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:30.979968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:8bb6cdb2627488d519412f81fa0ec971d1c2943a56f40000d26078bfb8696d90

Observation 6e298221-0052-4e78-bc0d-20742c6049ab · outbound

This paper cites an unresolved cited work.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-27T10:49:02.970934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:0f70a7f4d1750580cd6d0218f1889090173d624647f334df7b09f915d872fda6

Pith citing papers

Observation 03c64f5c-1682-4b91-a664-c94a02a0bffa · inbound

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures cites this paper.

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T00:25:08.066021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:25:08.066021Z digest=sha256:4efdff6fd8f0679c1355d37489730854493621b2da44df5ff11513def530685e