Pith. sign in

Paper Citation Record · LEDGER

Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2408.13006.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.13006 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:05:25.799258Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c736d4fb-3f19-43d8-bef9-b167f1e36ba6 · inbound

Towards an AI co-scientist cites this paper.

Towards an AI co-scientist Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:02:44.404608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T13:02:43.571234Z digest=sha256:bd0acb2e87b46c76fa045ba6e1179d20f18c0d3af5ea01c68932a8411410e6a8

Observation 57ba14f9-bb84-44fb-b442-9d6f7946c53a · inbound

Large Language Models for Predictive Analysis: How Far Are They? cites this paper.

Large Language Models for Predictive Analysis: How Far Are They? Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:05:25.799258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:05:25.799258Z digest=sha256:b3be55d22b82e036897ad0cac291ba0d406c7af9e9f05e834fc4922bae5f78a4

Observation b472d78c-92ea-4f16-8e06-1c1a56c278ee · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.693251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.693251Z digest=sha256:1cdf0544b41ffd7c76a04e7ae766102af14109f7b6e25ce168a22fa6ca030052

Observation 7c9c9731-28a4-45d0-914a-91332507933e · inbound

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance cites this paper.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.386305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.386305Z digest=sha256:a36252436bcf307c04ed446585bc937d0a791156f7e7d4f726b8e243b4842515

Observation ef315790-9b74-4f37-80c1-6a150431efaa · inbound

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation cites this paper.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.584381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:08.584381Z digest=sha256:d031b738a0c44d35893a8635a15300c1608ab2dd4b4abc25c6eb9a65ee02cab7

Observation 60835722-30b9-4029-95ec-185286065f5e · inbound

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability cites this paper.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.601561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.601561Z digest=sha256:25d40001f97827477b67ea51521f0cff7763c65cd3668093ccb1ad38bb9de574

Observation da944267-1197-4876-a579-307bef239571 · inbound

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing cites this paper.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.843092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.843092Z digest=sha256:ab336650037c0b36fea0846317d957e291fd0e6639ecb388990fb080f422268b

Observation 31fef8a8-1a7f-49d7-b3b8-a8bf0d604315 · inbound

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing cites this paper.

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T19:12:05.596485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:12:05.596485Z digest=sha256:bbe849f2a2273b025a017ed594b0208ff6ed94eca014a26e13b5612705fc4227

Observation 1dabb476-796a-40e7-a1a1-95629d87f9c2 · inbound

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs cites this paper.

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:14.617849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:14.617849Z digest=sha256:d6abaf72ebfb5693312b658d68fc53b6ab5b7d33c3ac5aa311204615f66165f1

Observation dd0bb244-766b-40a7-bb26-8341a1390907 · inbound

AI Propaganda factories with language models cites this paper.

AI Propaganda factories with language models Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T15:21:32.643649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:21:32.643649Z digest=sha256:1cfced33d805417ceec53a504982c5f62b50d21d596f2edd6a2c4603dcc04105

Observation 7c465ca9-2c19-4f63-a87d-aee427466f19 · inbound

IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations cites this paper.

IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T11:26:27.543864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:26:27.543864Z digest=sha256:0ec7d3b9f988b0f89bc4683a1c3ce36499cd5703b474c4f30629bfb96b2bc0de

Observation aeb117fe-1c82-4dba-8867-5cee665d83b0 · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-13T15:55:53.399860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:55:53.399860Z digest=sha256:3b2e0eb8f492d8a78732b1c0c8a982ca20d66b38cbc36f32d26feea5619cfa70

Observation d49d3437-8437-407b-8c8d-2b048b58bf33 · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T17:09:18.168845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:09:18.168845Z digest=sha256:db48a987a889001496d7cf69c9caefd9ad7a0e23a3099a72e11f627f274ac4d0

Observation d1d85efb-00b7-4abd-8b08-b003eaafd557 · inbound

LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection cites this paper.

LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:45:48.017272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:37:19.360378Z digest=sha256:852df41915e9596df491eee9e4bf7af705a6a792ddabb45b3bba04ce940b3234

Observation cfcb7436-1bf7-4185-b89a-df84914418af · inbound

Iterative Finetuning is Mostly Idempotent cites this paper.

Iterative Finetuning is Mostly Idempotent Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:44.631356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T19:04:36.210200Z digest=sha256:80a2345a4ccf6f71f4641ec5397c5343aea72eec281134cf1a19cbd3788424ee

Observation f2209723-09be-42a7-9623-9f00e582529d · inbound

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration cites this paper.

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:44:40.790933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T13:38:51.424919Z digest=sha256:56374e6f9a98c9a16d999a29edcd976997f46908bc06a950bb839d8f884c9883

Observation 7133626b-ec70-40b2-b733-4667d82b46c1 · inbound

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators cites this paper.

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-06-27T21:41:18.486232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T21:39:26.268337Z digest=sha256:aa6c5f238c68c04717bf0faf06f197096d063a4f68ad35ce832f0245c00d63e9

Observation 8ff3a2b6-4a2a-453b-8bed-405c8d29267d · inbound

Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment cites this paper.

Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T14:02:41.517301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:02:41.517301Z digest=sha256:c1a3995db4ba68ff1a03642fa271f67ccfb2fe064d77c3a38e9a054a8fee7737