Pith. sign in

Paper Citation Record · LEDGER

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution

As of 2 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2605.06125.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.06125 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T09:04:06.347354Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:47:01.373765Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact9
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c56448ec-fe67-4be4-b095-fed53ac09e85 · outbound

This paper cites TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:12.022865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:6ba845ac3ca4fd828c376a16fd55e85ee247e844354b52443159020829a87912

Observation 0325c06c-79f0-4450-b662-d01a85937647 · outbound

This paper cites SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:12.077714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:b589cfc9460f21bdb5e76d1fc0f208f649ce4ef80c4844dc06f859457b4b9358

Observation 95194f36-9e8e-441c-991a-63883f4a566a · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:26:12.004352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:ad5cb89e578fe9aa8034108a6271fa8de7d5ec850ef0b669dffbc06cc7a17eb6

Observation e5097a72-fd71-451d-b850-25b82c93f511 · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:93e96871b81b635974d42f992768b808ad0d9f7d34f566fc66f308da95a23419

Observation e46be76a-84d0-46be-87f5-0fc62ff52bef · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution GLM-5: from Vibe Coding to Agentic Engineering

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:26:12.010118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:6574a2374e4c06b2bb760d9927f39312ae0b8999144b525f16cb45bb559dcb89

Observation 36d55959-6cc0-4892-bcbb-70d17e6c3419 · outbound

This paper cites Just, R., Jalali, D., and Ernst, M.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution Just, R., Jalali, D., and Ernst, M

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:28:03.244773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:0f779e91f9f31773beb4fc8e0a908e46e66e6d75034b347e75dc10db44a48563

Observation c58597ae-8aeb-4413-b338-28b08539beb4 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution Kimi K2.5: Visual Agentic Intelligence

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:26:11.984427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:507fe095e6038607f4cfd825e618e205bbabbe9780947b28d6c68ee69e04fd2c

Observation 9ce0dc4d-c8a0-457e-9668-a9a46654f7be · outbound

This paper cites Qwen3 Technical Report.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution Qwen3 Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:26:12.040198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:c41805f111cf66061c648eadd687aa6e61f27b11570f21e862e7c493b4b46e64

Observation db750414-5a0e-476d-8032-6b5f0586f0f6 · outbound

This paper cites SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:26:12.030598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:eb543ed46300e6fe1084e5c6ffe31e570e8ca52fb15fe32c3dcc5e939796071e

Observation 911de558-d813-44b2-9a25-b6b6c46b441b · outbound

This paper cites Exotic Topological Phenomena in Chiral Superconducting States on Doped Quantum Spin Hall Insulators with Honeycomb Lattices.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution Exotic Topological Phenomena in Chiral Superconducting States on Doped Quantum Spin Hall Insulators with Honeycomb Lattices

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:11.976093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:5ce07d2289382af4f8be62426e597abd54eb31fbb6f65e07790d38ac70eb29d3

Observation 2ae7aeed-dde3-43fc-ad02-536464b3ff4d · outbound

This paper cites Advancing Code Coverage: Incorporating Program Analysis with Large Language Models.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution Advancing Code Coverage: Incorporating Program Analysis with Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:11.944425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:88e16f5ed693d072d0daab609414fe2252d9c2ca639fcc6a4f3beec7622cf76f

Observation caa121fa-d441-4354-8d44-8ccb4580c19a · outbound

This paper cites Unit test up- date through LLM-driven context collection and error- type-aware refinement.arXiv preprint arXiv:2509.24419.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution Unit test up- date through LLM-driven context collection and error- type-aware refinement.arXiv preprint arXiv:2509.24419

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:11.962909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:e4d882517eb6189bf9679f7a8089f84d8ed13e2ec3c04cc3f6a440c43b33a1b8

Pith citing papers

Observation 549c5b4b-52d0-4b48-9297-33de4b6f62c9 · inbound

MultiFixer: A Coordinator-Proposer Based Multi-Agent Framework For Fixing Multi-Hunk Bugs cites this paper.

MultiFixer: A Coordinator-Proposer Based Multi-Agent Framework For Fixing Multi-Hunk Bugs Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T12:47:01.373765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:47:01.373765Z digest=sha256:f1385072278a16c6029b7b2d2ad96d660c1fcc32933e3af2fc5c81084980b466