Pith. sign in

Paper Citation Record · LEDGER

Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2502.06193.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06193 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:22.560982Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:55:22.701914Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b0bf9b93-84a3-4696-8c9a-dc533fb2cb45 · inbound

Larger Is Not Always Better: Exploring Small Open-source Language Models in Logging Statement Generation cites this paper.

Larger Is Not Always Better: Exploring Small Open-source Language Models in Logging Statement Generation Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:22.560982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:22.560982Z digest=sha256:608d63ff97ce78036c6351c84bea560ec1ca2ed439704509f904da51b202e066

Observation ac31a24c-c423-4aa8-bf6f-cc421c3367c4 · inbound

CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review cites this paper.

CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:31.407675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:31.407675Z digest=sha256:2be81b5e63fb5be87c5cf64dcb84f48b4ebb135b51b66822cfa12c17af700466

Observation 1a0d2f36-7f99-4b90-863c-425c93e235a5 · inbound

Querying Large Automotive Software Models: Agentic vs. Direct LLM Approaches cites this paper.

Querying Large Automotive Software Models: Agentic vs. Direct LLM Approaches Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:17.022076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:17.022076Z digest=sha256:29f515e9e7060d292747bcbaed3b04eb9f4e512a1b79eec01d4c7492d17e85aa

Observation f548eb6d-8b09-4ea9-83a0-367b254dba3d · inbound

Spiritual-LLM : Gita Inspired Mental Health Therapy In the Era of LLMs cites this paper.

Spiritual-LLM : Gita Inspired Mental Health Therapy In the Era of LLMs Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:14.277116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:12:14.277116Z digest=sha256:46432060ff63af0ee4b69ce7509e23373abd914ad483ed0170db0011d8dafe91

Observation 9462f634-0751-458c-984c-38e6b37bddfb · inbound

CCISolver: End-to-End Detection and Repair of Method-Level Code-Comment Inconsistency cites this paper.

CCISolver: End-to-End Detection and Repair of Method-Level Code-Comment Inconsistency Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:55:22.740075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:55:22.623455Z digest=sha256:5cdef2bf06b4972163763057f29e0e6e415f30779bc9afd145df24f693c4304c