Pith. sign in

Paper Citation Record · LEDGER

CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2502.16614.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.16614 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:00:17.038881Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bb73fe5c-1553-4d57-80b8-eec45055522d · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.038881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.038881Z digest=sha256:32cddd4c5bcae9d6d7e7e240a07c8a7dd411d253029148aea01ee1a02dce8d2b

Observation a807d40b-dbf3-4e99-ae99-5d282ba09b7f · inbound

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization cites this paper.

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:21.260021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:14:21.260021Z digest=sha256:2cb7c4a8678d5fcb3a41fb1c59fd14c6ce78f6d11cd5f67663a2b0c721fde427

Observation ce05db4f-2cd0-4649-835a-3c41266d1bce · inbound

Rethinking Technology Stack Selection with AI Coding Proficiency cites this paper.

Rethinking Technology Stack Selection with AI Coding Proficiency CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T17:07:47.167602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:07:47.167602Z digest=sha256:c0413716cf5eb810d5d3883d824545caa7cf511948a043692561d328a4547b50

Observation 857e918f-d033-4ba3-9b06-efdc04d1fd06 · inbound

Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces cites this paper.

Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:37.874477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T18:44:27.685266Z digest=sha256:9c4ed0dae2989c45a9f6a17f092ffeaaa2ff618fd7599139d4c9f272f9156cb8

Observation b0335a26-87c7-4c0a-a678-558be01b53d2 · inbound

DeployBench: Benchmarking LLM Agents for Research Artifact Deployment cites this paper.

DeployBench: Benchmarking LLM Agents for Research Artifact Deployment CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T09:16:49.251792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T05:35:12.027697Z digest=sha256:a0a1f27e9f24d302490cfecf34011c509ead0d78be040fef7d376fa7e379e50b

Observation a9a87fbc-5af8-4b64-81d1-579e2c4b1ae8 · inbound

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces cites this paper.

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T22:31:21.480801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T22:22:52.690010Z digest=sha256:9ab5093916816f891350982a912057bef807dfa427b1805a9a11d373a52076f8

Observation 36ed1f45-04aa-4eeb-b464-4e68cab1547a · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models

Reference 173

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:17:24.923463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:4b092281b93110583c34a007380be46d46a940fff7cd050d5716597760cd991d

Observation d4c6c586-f48f-4110-a161-fc8569ddce59 · inbound

Code Monitor Red Teaming for Public-Test-Passing Code cites this paper.

Code Monitor Red Teaming for Public-Test-Passing Code CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T09:12:27.327027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:12:27.327027Z digest=sha256:5f436a06943332e3b5ef3280646a4b82fd2a7c1640221aa65ea621c35373e4fa