Pith. sign in

Paper Citation Record · LEDGER

Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

As of 6 August 2026, this Paper Citation Record lists 2 of 2 outbound references and 7 inbound Pith citation observations for arXiv:2605.27922.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.27922 v1

Coverage vector

measured 2 of 2 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T10:51:01.483339Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T23:26:36.859251Z

Reference resolution

2 of 2 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7c08f4d8-66a1-4162-878d-5249dfb3fea6 · outbound

This paper cites Holistic agent leaderboard: The missing infrastructure for ai agent evaluation.

Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows Holistic agent leaderboard: The missing infrastructure for ai agent evaluation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:53:26.437062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T12:52:19.145898Z digest=sha256:dec610d91683380864a513c58c0a0da3670b864f39a66d7e2c52ce0be2d07b5c

Observation 0a41e618-6075-49b8-9be6-e83fe4f8838d · outbound

This paper cites Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents.

Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T12:53:26.226692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:79deb62a2dfdb9d846936a0fe3abbae59851db8f8cd8e9afc311bd146c60baa5

Pith citing papers

Observation a0b0d40a-e0c0-4065-9be1-386ae646d9a6 · inbound

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions cites this paper.

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:36:29.209900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T09:56:36.860369Z digest=sha256:83ff33e3de4285c973ff5f629489f72a07bc1128f8bf58736bd7df6211d168a3

Observation fad251e8-db19-4c39-b790-b905800a307f · inbound

Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI cites this paper.

Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T06:28:52.906906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:28:52.906906Z digest=sha256:493db482d70324cd2c174f70ab5f628426faa0172ea245468621919ab17c26d6

Observation 920c8117-e741-4813-8d5e-253d5b6cb730 · inbound

The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI cites this paper.

The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-09T23:26:36.861144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T23:24:07.130101Z digest=sha256:ebf23d110861203e7b22063d24677684f0e2fb55d821d5e42950aecf0959eb0a

Observation 36e6465c-0460-4ded-b775-0bc2ff6f29b0 · inbound

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading cites this paper.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T05:28:45.311405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:28:45.311405Z digest=sha256:c4f6a822d1ebc297d494641732561e6f2bdb81fdf1b55ec9a1ccaa5d45ab43ac

Observation 1e3afeb1-d4eb-4dbd-85be-d2e879f716e7 · inbound

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading cites this paper.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:c3018a335ec247bb4e1c95a35b54b60b62a9fdc992fbc9981ad0392ecc2d38fa

Observation 96eb98cb-443f-4d8b-826b-3e1da8ddcb0d · inbound

Two Confounds in Cross-Model Value Comparison: Response Determinism and the Access Harness cites this paper.

Two Confounds in Cross-Model Value Comparison: Response Determinism and the Access Harness Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T13:32:31.523845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:32:31.523845Z digest=sha256:4d0c89c35d60348c70dc42f31dc921148e7511bdabcf70fd3e65c8e61be588c6

Observation 09084219-0f7d-47c0-9043-ba7a73ebd912 · inbound

Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results cites this paper.

Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-01T10:51:01.483339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:51:01.483339Z digest=sha256:64d04e7bd498247189b41ffb5cd1aad33b437a8b807f3795735ffb93385b66f4