Pith. sign in

Paper Citation Record · LEDGER

ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2502.05352.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05352 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T07:41:24.219547Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ddff6922-1cf3-4055-83be-510cd414c2ca · inbound

LLMs Corrupt Your Documents When You Delegate cites this paper.

LLMs Corrupt Your Documents When You Delegate ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:48:47.626847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T09:47:21.966292Z digest=sha256:0b7b5e1c745666775ef6c053f1d23c9efba069a7b454651fae802c5ad89deacb

Observation 0247fec9-b64b-4ff5-b8d1-8863f7809c47 · inbound

Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery cites this paper.

Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-05-08T20:04:06.598794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T17:28:41.217810Z digest=sha256:5d87565e6fbe00c44336a59c42b41123773f962369ef244d542c7a570293f537

Observation f603da73-d2fd-44f9-b490-51cab57d9fb7 · inbound

Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery cites this paper.

Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:13:52.344837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T00:13:07.546472Z digest=sha256:e1b96ee80da7bd97ee6a067f4bc73d63dace2593bac4317ec67c0d34a1610cfc

Observation b341aca9-665d-4d30-a077-6ec02259c301 · inbound

From Assistance to Agency: Rethinking Autonomy and Control in CI/CD Pipelines cites this paper.

From Assistance to Agency: Rethinking Autonomy and Control in CI/CD Pipelines ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:57.170897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:06:14.217588Z digest=sha256:bf82473e5b0551b538c7a3fa484d2f6d0fb48e8ec837efe5af170c0c84b2302e

Observation df8a5b55-37bf-49a0-8ef3-7ecabce42b99 · inbound

Runtime-Structured Task Decomposition for Agentic Coding Systems cites this paper.

Runtime-Structured Task Decomposition for Agentic Coding Systems ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.175763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T14:46:37.166185Z digest=sha256:032fb4c7bfe153ce2ad202dcb23ed69a33c733e2f708203e452fc6d612943433

Observation d0025bab-9094-4305-adcc-ad8f611abf7d · inbound

Auditable Graph-Guided Root Cause Analysis for Kubernetes Incidents cites this paper.

Auditable Graph-Guided Root Cause Analysis for Kubernetes Incidents ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:27:28.084149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T18:10:31.747998Z digest=sha256:365676807935dd44128c66fd7c5f40273cee1fa5a27e9a09a686c4d799d72295

Observation ecaeffc4-54e8-476c-8a4e-e48ca3bdc2a5 · inbound

Pooled Leaderboards Hide System-Specific Winners: A Reporting-Protocol Audit of Offline Root-Cause Analysis Benchmarks cites this paper.

Pooled Leaderboards Hide System-Specific Winners: A Reporting-Protocol Audit of Offline Root-Cause Analysis Benchmarks ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:54:21.844748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T07:52:33.302523Z digest=sha256:5d2e3ad9fb5a13ea5409df053957fe42e1bba79f846705cf92999e0db63c78c2

Observation 7e62ae76-ef13-4998-bb89-f97064594ab4 · inbound

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents cites this paper.

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T12:47:54.895493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:47:54.895493Z digest=sha256:3aecebd5e3585b574441080fee6a81aba938abbbf952b9036242d8226b0a5b21

Observation 4886e0b6-7834-48f1-94b6-916a93d33fd3 · inbound

Beyond Component Testing: Validating Agentic AI Systems cites this paper.

Beyond Component Testing: Validating Agentic AI Systems ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T07:41:24.219547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T07:41:24.219547Z digest=sha256:9b2779d6c98a90ba16bb491a0636c6d8fb4d2061d10f9dff32f4b3d629b094e8