Pith. sign in

Paper Citation Record · LEDGER

ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2502.05352.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05352 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T07:41:24.219547Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ddff6922-1cf3-4055-83be-510cd414c2ca · inbound

LLMs Corrupt Your Documents When You Delegate cites this paper.

LLMs Corrupt Your Documents When You Delegate ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:48:47.626847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T09:47:21.966292Z digest=sha256:a3842e63cc4664be09ab725222db96706320434e676145ac5072d60851e1bf36

Observation 0247fec9-b64b-4ff5-b8d1-8863f7809c47 · inbound

Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery cites this paper.

Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-05-08T20:04:06.598794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T17:28:41.217810Z digest=sha256:80b4c210f783ab793694f95fca6b11e9b35eb3e3a497dabfb7942873fecf3675

Observation f603da73-d2fd-44f9-b490-51cab57d9fb7 · inbound

Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery cites this paper.

Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:13:52.344837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T00:13:07.546472Z digest=sha256:68b22a282d5263c9e8b3e2b924ab5b5e18f6c40979ba2698e5407b3dfb266327

Observation b341aca9-665d-4d30-a077-6ec02259c301 · inbound

From Assistance to Agency: Rethinking Autonomy and Control in CI/CD Pipelines cites this paper.

From Assistance to Agency: Rethinking Autonomy and Control in CI/CD Pipelines ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:57.170897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:06:14.217588Z digest=sha256:ab689ba3eaf5648f4897324da022eeef242b50bbb19f3a75f4180766ad4d178f

Observation df8a5b55-37bf-49a0-8ef3-7ecabce42b99 · inbound

Runtime-Structured Task Decomposition for Agentic Coding Systems cites this paper.

Runtime-Structured Task Decomposition for Agentic Coding Systems ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.175763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T14:46:37.166185Z digest=sha256:61585cf8d7bf9761bbbd0e9a0ccf3861969b9aa4e1a8ff185bc4b20d40509219

Observation d0025bab-9094-4305-adcc-ad8f611abf7d · inbound

Auditable Graph-Guided Root Cause Analysis for Kubernetes Incidents cites this paper.

Auditable Graph-Guided Root Cause Analysis for Kubernetes Incidents ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:27:28.084149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T18:10:31.747998Z digest=sha256:345e4bbd08f4d4c6a680600ba84267a7006eb3c0008bd0a662618ba9f6e1a17d

Observation ecaeffc4-54e8-476c-8a4e-e48ca3bdc2a5 · inbound

Pooled Leaderboards Hide System-Specific Winners: A Reporting-Protocol Audit of Offline Root-Cause Analysis Benchmarks cites this paper.

Pooled Leaderboards Hide System-Specific Winners: A Reporting-Protocol Audit of Offline Root-Cause Analysis Benchmarks ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:54:21.844748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T07:52:33.302523Z digest=sha256:0e5225ef6510f4faf41c21e6117fce2c2bd7333f2ed3e69024d4e268c78ceb88

Observation 7e62ae76-ef13-4998-bb89-f97064594ab4 · inbound

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents cites this paper.

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T12:47:54.895493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:47:54.895493Z digest=sha256:2d5b7e75404acf2152daed7fed6361cc978600ffeecd5f8668935bcc3d31fac3

Observation 4886e0b6-7834-48f1-94b6-916a93d33fd3 · inbound

Beyond Component Testing: Validating Agentic AI Systems cites this paper.

Beyond Component Testing: Validating Agentic AI Systems ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T07:41:24.219547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T07:41:24.219547Z digest=sha256:cb287d8f7ccc10ae0fef02d91beb353c854d6a0f19a5f35ca5fff4b5f4eaca78