Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T03:37:07.841385Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 100 inbound Pith citation observations for arXiv:2601.11868.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T03:37:07.841385Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T07:17:13.420563Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
20 of 20 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation ba6d4cc9-f1fd-4bd6-8dcc-cb5f6381aad1 · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces TextArena
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cd9c0420-6168-4263-b307-9dde5ac8dbb6 · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces VisualWebArena: Evaluating multimodal agents on realistic visual web tasks
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b485e30e-fadf-47a8-bb07-efe7236bdf3b · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces gpt-oss-120b & gpt-oss-20b Model Card
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 33ee05cc-1213-40ed-b33c-c6a79eff3c2c · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces K., Krupke, D., Kidger, P., Sajed, T., Stellato, B., Park, J., et al
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 66ecafb5-6e96-4e31-a056-9461b2744708 · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 95239b1d-5d3b-44fe-8c52-63b01aefd6d2 · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces 18 APPENDIXTABLE OFCONTENTS A Detailed Results 21 A.1 Comprehensive Results
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b7a596a1-1c32-4bbe-8ce8-93c5f9ed1886 · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e3bb58bb-4a28-4859-8580-3b1f89c34848 · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5f12f90b-f0f7-4a08-a6f0-9400bed77132 · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1ea20868-973d-44d0-b83f-5744ccd0f5de · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces outcome"
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 704a83f3-364a-48de-bd1a-fee2d68af6f1 · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 18234bb5-8c76-4752-91b3-3aa121ccce72 · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Independent inspection of the released tasks con- firmed that they are well-specified and largely free of ambiguity or underspecification
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 69736d52-131f-43c6-8a5e-72c1743023fe · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces held-out
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8420948f-997f-420e-8658-33e85060d913 · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 79ae3ee4-b721-4884-a081-76231d0e48a3 · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces We’re looking for indications that running the command has failed - carefully analyse the outputs to determine this
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b78106c3-ece6-4f01-8d47-5bb24889ec82 · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df70fb1f-1c80-445d-b7c3-83b63bab7be6 · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 60003a15-e549-4f65-bdbe-2021236c8a53 · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces You are about to hand off your work to another agent. Please summarize what you’ve done so far
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c92b3ab6-065e-40ca-9339-c9cd25485b5a · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 75a8499d-5053-419f-8936-085fa367be89 · outbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Results: X Y Z
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bc2c7d59-08bb-43f3-8808-9fbaa0217091 · inbound
SERA: Soft-Verified Efficient Repository Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cadcd68-2d55-4d6e-ac8d-6a4aa8a99808 · inbound
Kimi K2.5: Visual Agentic Intelligence Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8d070eb9-7a96-4dbd-8588-290400dabc58 · inbound
Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c2502f82-f2f0-4f31-bce7-5ad70f6b8715 · inbound
VeRO: A Harness for Agents to Optimize Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3f66d6e5-9cbe-41ef-9b2e-302297c6924d · inbound
VeRO: A Harness for Agents to Optimize Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e43ae4b7-e82f-472f-af09-1d853d86e373 · inbound
Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c361e55c-5065-4bb8-bfcb-60f76da75665 · inbound
Effective Strategies for Asynchronous Software Engineering Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50b1f0e6-6b95-4ac1-a3fe-226c9e4e69ea · inbound
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c9e2e6ae-129b-440b-8428-c68864acdeec · inbound
Meta-Harness: End-to-End Optimization of Model Harnesses Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 60c625e6-06e4-4417-aa07-988588164588 · inbound
Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b16f0811-0ba8-4acc-89f3-22f44b099f5d · inbound
TRACE: Capability-Targeted Agentic Training Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2c22284f-4079-4a85-aa42-a9581bc39753 · inbound
AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a383907e-af74-40bd-9195-d767b431b951 · inbound
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9d40e8b6-cfd2-4416-a969-90a908bbc508 · inbound
Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4818b396-6ccc-4d60-9eec-fce92bdba149 · inbound
Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecd9eaa4-0a65-46c0-a623-229feafe242a · inbound
How Much Heavy Lifting Can an Agent Harness Do?: Measuring the LLM's Residual Role in a Planning Agent Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 01afe455-1add-420e-9124-63c932f63770 · inbound
Prediction Arena: Benchmarking AI Models on Real-World Prediction Markets Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f81e0b6f-b431-4ec2-808e-eeb124a87ee5 · inbound
COMPOSITE-Stem Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d1cb947c-175e-4db4-9988-9380e13e9586 · inbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 31d7448c-c4a4-4761-b50f-089bd20710e4 · inbound
From Context to Rules: Toward Unified Detection Rule Generation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 24ac4b31-03b1-4834-8f89-731d838722dc · inbound
From Translation to Superset: Benchmark-Driven Evolution of a Production AI Agent from Rust to Python Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b4eb09ff-74d5-4d1f-a6e4-066780d33ace · inbound
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9b070afe-5685-4a0c-87fd-f39ebfa076c0 · inbound
Towards Long-horizon Agentic Multimodal Search Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 342b5f46-0949-417f-b797-6b2d68b7ad42 · inbound
Exploration and Exploitation Errors Are Measurable for Language Model Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ce1cea9c-d44b-491e-86d8-291b39ac4e5b · inbound
Don't Let AI Agents YOLO Your Files: Shifting Information and Control to Filesystems for Agent Safety and Autonomy Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d6db8f4c-b25d-40c8-9fc5-d247cdfbb397 · inbound
HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e22b2236-bded-4744-b39e-beffc944a32f · inbound
KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ebf058fe-d98b-4a30-a5a3-ba495e4e3a9a · inbound
BranchBench: Aligning Database Branching with Agentic Demands Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 222dc975-e5bb-4afa-974d-e8eea22f129d · inbound
SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b3620eee-499e-42a7-a8df-cab7d7b76c02 · inbound
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 73a253dd-f2fd-451e-b667-631dd09377f5 · inbound
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a4525fe0-769f-415d-9c41-d2dbc0339314 · inbound
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8c5651dc-ed4c-4bbb-8999-66f412d579d6 · inbound
Toward Scalable Terminal Task Synthesis via Skill Graphs Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96afacb2-6ed8-4f57-a9e4-1b1a3d06d911 · inbound
Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a92fa150-3bf8-4f25-b86a-c31343bdc321 · inbound
Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0fac8e12-f199-403e-8617-2cb5f533c068 · inbound
MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2f684907-36ef-4aff-a4df-fb2158e37a91 · inbound
Heterogeneous Scientific Foundation Model Collaboration Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b7cf7a6-cd43-4d3a-949f-0897f251da6e · inbound
What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9546a39b-4954-431b-9a4a-4abc9fc99d66 · inbound
Crab: A Semantics-Aware Checkpoint/Restore Runtime for Agent Sandboxes Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6bdf56a8-1aed-4386-ae4d-fb2556d38d18 · inbound
LLM-Oriented Information Retrieval: A Denoising-First Perspective Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 126
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1c02537c-702a-425d-9fc1-fb4fe28c57a1 · inbound
LLM-Oriented Information Retrieval: A Denoising-First Perspective Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 961c9605-9309-40bd-9df3-2f434e68e2d8 · inbound
BUILD-AND-FIND: An Effort-Aware Protocol for Evaluating Agent-Managed Codebases Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6c3523ec-d2b4-49d8-baf1-3470940b2b51 · inbound
PrefixGuard: From LLM-Agent Traces to Online Failure-Warning Monitors Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0ebd8258-2635-42b0-acfa-36a76a63ddb0 · inbound
TeamBench: Evaluating Agent Coordination under Enforced Role Separation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bb9829ac-ef4d-47a2-961c-d4ea2a3aded4 · inbound
SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 61cd7e8a-a615-4554-b1d0-b48146f8d843 · inbound
SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation decc0cb2-5d3a-4d38-bb4c-c0ce013a4088 · inbound
SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aebe349-74e9-4a5e-8932-3b0870b4ed46 · inbound
Learning CLI Agents with Structured Action Credit under Selective Observation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9460b06a-2718-4a8d-824e-c7412ab0b561 · inbound
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dcafdf38-1693-42ab-86cf-5468dfba4fc0 · inbound
MDGYM: Benchmarking AI Agents on Molecular Simulations Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df52b1ab-c41e-4b5a-a125-ec598d19c1a2 · inbound
LLM Agents Already Know When to Call Tools -- Even Without Reasoning Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f2f9c64f-d445-40d5-a80f-22e192a3ff2b · inbound
LLM Agents Already Know When to Call Tools -- Even Without Reasoning Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f3f32258-9907-456a-953e-b8350f686e5c · inbound
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e730d03d-012e-46db-ba61-4682e42e3a8a · inbound
Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bf81f40f-85c9-4ef1-b7a0-a9afd75ccfb4 · inbound
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6ff4237c-23a0-432b-b0dc-b69a0e5b4df1 · inbound
Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fca59e70-5b37-42a0-81e0-55231c125552 · inbound
Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 51e3bf8a-4d6a-40bd-827e-3b03fabf1c24 · inbound
MMTB: Evaluating Terminal Agents on Multimedia-File Tasks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a5c385cf-7725-43df-88d9-56073daa5181 · inbound
gwBenchmarks: Stress-Testing LLM Agents on High-Precision Gravitational Wave Astronomy Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c32812f1-7fb8-42ee-940e-4ce580779882 · inbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 503b79ae-787f-49fa-83da-4f68cb2e5327 · inbound
AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 15c81c16-6fee-4891-adf5-99f8d6bc80d6 · inbound
PREPING: Building Agent Memory without Tasks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 309e8e6a-3dae-44f3-8ced-fbfd0fda3b69 · inbound
Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4fc34411-3e50-43d2-ba18-160b02c5230f · inbound
CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aed8048e-41c0-4af3-892a-79f4cf2aa89d · inbound
CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 78cfac98-5e6c-4228-a650-34cfaf220276 · inbound
ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6ea8b4e0-dd3e-471f-9a61-85c3ec0be6e3 · inbound
ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e219b491-3529-4255-91c5-73f8bd914e35 · inbound
Do Coding Agents Understand Least-Privilege Authorization? Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2f93dc2e-a6e9-43c4-845e-e4d0be8ab33a · inbound
Orchard: An Open-Source Agentic Modeling Framework Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 88082ec2-1c15-4900-a813-9175ecfadaf6 · inbound
Orchard: An Open-Source Agentic Modeling Framework Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27abc1ab-56ef-4fc6-9416-469e0595125f · inbound
The Scaling Laws of Skills in LLM Agent Systems Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a69b9c14-c1b8-4e96-acff-34dcef2fff66 · inbound
R2V Agent: Teaching SLMs When to Ask for Help Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 75a26867-59a5-44f7-9907-2686bbfb54e2 · inbound
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 715c29b0-5284-465f-88ec-2763e48ffdd8 · inbound
Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 45f591ad-0cee-4f48-871a-0c59c6c30f26 · inbound
Responsible Agentic AI Requires Explicit Provenance Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 532443f7-28e3-4935-ac72-dc09c9cf633a · inbound
WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 93d737d9-6e7b-42b8-aaeb-c649ceb56585 · inbound
WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ce1150a4-67d4-4813-9d06-6f50130693db · inbound
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e97fe467-d648-43b5-b617-064f30d817af · inbound
HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3cc47777-f3ce-4c2a-acd6-9db4470922e5 · inbound
Open-World Evaluations for Measuring Frontier AI Capabilities Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c3805acd-f5f9-4626-8013-2009cb555e79 · inbound
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9603e92f-7fe6-420b-985b-9c7047c38e74 · inbound
Terminal-World: Scaling Terminal-Agent Environments via Agent Skills Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0548592f-515c-49c2-a65e-efb64215ce0e · inbound
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c573288-681b-4aeb-a22e-3feb1f2a2278 · inbound
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 68cb9568-1160-4b29-bea1-9285dd28addd · inbound
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5ee85988-18d5-4098-a1f8-a56ac7e999f1 · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation de70236e-354d-44c0-a4b1-0b1eeace607b · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 12e2610c-841b-4ae3-a2c5-0ffe5acb0b16 · inbound
Stop Comparing LLM Agents Without Disclosing the Harness Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c6e275ad-ec3e-4325-90da-d6c1b7f0fded · inbound
EvoCode-Bench: Evaluating Coding Agents in Multi-Turn Iterative Interactions Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 841dbef7-3a3c-43cb-ab9a-74663fcee331 · inbound
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation deefa0d1-ddb4-4fb0-8e1c-bc03ee987e64 · inbound
DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 28ff39d4-3f31-448f-8496-ced88e58a6ae · inbound
Heimdall: Formally Verified Automated Migration of Legacy eBPF Programs to Rust Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2a338a80-4f3c-4d43-95ad-779cfd6205db · inbound
From Model Scaling to System Scaling: Scaling the Harness in Agentic AI Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fe3a6ea6-693b-4610-a129-dfce44b43809 · inbound
Agentic AI Workload Characteristics Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e274befe-b446-46bb-b4a5-9234ca378836 · inbound
Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bbb67e4c-f862-46d1-a376-9406eedf60ca · inbound
LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 54a1d8d7-304f-474f-b75a-e6fb64032235 · inbound
TraceGraph: Shared Decision Landscapes for Diagnosing and Improving Agent Trajectories Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a0b672bc-4ad4-429b-8d7d-b5e3125865f7 · inbound
Stateful Online Monitoring Catches Distributed Agent Attacks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c1ef4e1d-6e4b-4cb0-b36a-03adb6528df3 · inbound
ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1404f36a-9048-437a-aee0-5f5b65eb72eb · inbound
Sakura: An Approach for Generating Complex Tests from Natural Language Test Descriptions Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.