Pith. sign in

Paper Citation Record · LEDGER

CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2408.10718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.10718 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:50:50.690353Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T17:35:44.171248Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b5c19431-5940-4394-a6e7-465c85ba4efa · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

Reference 219

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:44.174084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:eddd83f919e822cacf6f548968faff9a3154c7c712ce081f45ee26dff9805358

Observation 6d27ed5e-6df0-4916-8a91-08e3c4982c17 · inbound

Evaluate-and-Purify: Fortifying Code Language Models Against Adversarial Attacks Using LLM-as-a-Judge cites this paper.

Evaluate-and-Purify: Fortifying Code Language Models Against Adversarial Attacks Using LLM-as-a-Judge CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:50:50.690353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:50:50.690353Z digest=sha256:c987cb355d070b108ff60b6f0b79c946bd04352c315c9f1168319bc94de6f006

Observation c59a42ba-7e1d-4028-9e5f-267ad9b6eeb7 · inbound

Is Your Automated Software Engineer Trustworthy? cites this paper.

Is Your Automated Software Engineer Trustworthy? CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:06:53.978307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:06:53.978307Z digest=sha256:64d4a985a3613479a8cbf3c5a568219b06fe460403ec6b45d9427fa24fd303c3

Observation 1db605c3-6c3c-4b88-be35-b78c37460583 · inbound

ReCatcher: Towards LLMs Regression Testing for Code Generation cites this paper.

ReCatcher: Towards LLMs Regression Testing for Code Generation CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T17:57:05.981261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:57:05.981261Z digest=sha256:8b828ffdae9564bbdb624322758763e6583f8d0b43b07898e1a743e5c4792dae

Observation 195b9ce0-c22c-4dfb-a4ab-6043b7c4c515 · inbound

Enhancing Large Language Models with Retrieval Augmented Generation for Software Testing and Inspection Automation cites this paper.

Enhancing Large Language Models with Retrieval Augmented Generation for Software Testing and Inspection Automation CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:44:38.021334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T10:40:30.734804Z digest=sha256:837346628ec934f0f5235d45e2a27e056d713acfd23a1d242a3b2dfcb40a800a

Observation 8d28c040-a207-4ebe-b0be-c058de436af6 · inbound

Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering cites this paper.

Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:32:00.279793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T07:29:03.994957Z digest=sha256:66f089cb767c6619b37f70d3616be6a5d52a37d82ff03af8a1e5314203cc226f

Observation 8b33dec6-9523-40c2-b587-d803d99e5a26 · inbound

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents cites this paper.

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:26.012242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T01:12:37.970638Z digest=sha256:fa3634ca5eddbdd1247b6ec65f574f74fefebbc34d4c52c7b2f985117646b9da