Pith. sign in

Paper Citation Record · LEDGER

JudgeBench: A Benchmark for Evaluating LLM-based Judges

As of 7 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 39 inbound Pith citation observations for arXiv:2410.12784.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.12784 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T01:24:02.052803Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 39 of 39 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:40:03.461298Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T17:09:58.526716Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22e41b80-1d6d-44ee-834b-0e09ebfc8657 · outbound

This paper cites 10/10”, “Neither A nor B.

JudgeBench: A Benchmark for Evaluating LLM-based Judges 10/10”, “Neither A nor B

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:24:02.075449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:24:02.052803Z digest=sha256:68385ea7c42c2b1e52a304323cb31eeeeb32baf130e5a467055056c8bc8cd4ce

Observation c203191a-bd16-4eb5-b2a9-8dee1cb61036 · outbound

This paper cites knowledge.

JudgeBench: A Benchmark for Evaluating LLM-based Judges knowledge

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:24:02.082283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:24:02.052803Z digest=sha256:52238a79189821a07339759b9fddac96c8468b181a7f27aa0fae7bfb8a721c62

Observation 6a66ec54-bfd4-4d96-bc51-1dbb4b058d16 · outbound

This paper cites Each side’s lateral pterygoid has a different function during this movement.

JudgeBench: A Benchmark for Evaluating LLM-based Judges Each side’s lateral pterygoid has a different function during this movement

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:24:02.087724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:24:02.052803Z digest=sha256:2b4448d8450d9ae64c9a1fc69d9f91ebc4f66e5c49b2b93e0015ed150f30cf14

Observation dadbf0f8-afce-4ed4-b5d6-126c97d5ffb5 · outbound

This paper cites an unresolved cited work.

JudgeBench: A Benchmark for Evaluating LLM-based Judges Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:24:02.091561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:24:02.052803Z digest=sha256:e6fe206448cc902ebde8e6db11fd807020ff75602528143dc357c730a02c5caf

Observation 4406d922-8fec-4c77-8c71-8752311c2bda · outbound

This paper cites Output (a).

JudgeBench: A Benchmark for Evaluating LLM-based Judges Output (a)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:24:02.095145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:24:02.052803Z digest=sha256:6a8bca86a279648f5ae2cda06c456d15d9b232ada3e498417147cc0d6df51f9a

Observation 49b9ea51-6e31-40e1-94c3-0c07e9563525 · outbound

This paper cites an unresolved cited work.

JudgeBench: A Benchmark for Evaluating LLM-based Judges Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:24:02.100089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:24:02.052803Z digest=sha256:3942b30721446c61e2a4d2e4f6c4ca6684c54d639cba85862562fd1bbe6a6079

Observation 6fe14601-6ef7-4601-bff3-83033f4981c2 · outbound

This paper cites an unresolved cited work.

JudgeBench: A Benchmark for Evaluating LLM-based Judges Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:24:02.104086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:24:02.052803Z digest=sha256:3c83bba8c06dc367371d484dbee85577547943725ded3ef5185144c32d63c62e

Observation c72140b1-2a06-4ac1-953b-f1c43386206c · outbound

This paper cites an unresolved cited work.

JudgeBench: A Benchmark for Evaluating LLM-based Judges Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:24:02.108077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:24:02.052803Z digest=sha256:4b796e211d9c361436ffc8a1f833f362d3e843bf64cb30bb5fd59b778b7b0465

Observation 8d1b2595-9978-48ac-b65e-3b00f315782b · outbound

This paper cites an unresolved cited work.

JudgeBench: A Benchmark for Evaluating LLM-based Judges Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:24:02.112140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:24:02.052803Z digest=sha256:6e4a12e81d12185525f0102df3400c6e733362b0f8b19c2334ba9cc29f3db35e

Observation 80a5f0f9-fdb1-45e0-840e-81cb9ebae46c · outbound

This paper cites My final verdict is tie: [[A=B]].

JudgeBench: A Benchmark for Evaluating LLM-based Judges My final verdict is tie: [[A=B]]

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:24:02.115722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:24:02.052803Z digest=sha256:3905469294bc93874ed5c090fc02d61e353d9cfacfabacb8306d4385671d5f31

Observation 19303786-670b-4a39-9c58-f6e3fbef159d · outbound

This paper cites an unresolved cited work.

JudgeBench: A Benchmark for Evaluating LLM-based Judges Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:24:02.119676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:24:02.052803Z digest=sha256:c1844791f6dfdd562ce72778fedc5707fa223f91bc3ca80b2f715be292a192a6

Observation b9b1432f-66fc-4bc3-a9fe-90c56b394e53 · outbound

This paper cites You should refer to the score rubric.

JudgeBench: A Benchmark for Evaluating LLM-based Judges You should refer to the score rubric

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:24:02.124472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:24:02.052803Z digest=sha256:84223e965ebdf629d4027ed6946841cd0dd9d0c80d3c02018926440e7beb15c5

Observation e43529e3-e0bb-4cd2-9852-7707b7a9caef · outbound

This paper cites (write a feedback for criteria) [RESULT] (A or B).

JudgeBench: A Benchmark for Evaluating LLM-based Judges (write a feedback for criteria) [RESULT] (A or B)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:24:02.128528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:24:02.052803Z digest=sha256:93c9a84c969388e03395cdd43d74b5deac4d0da88d74485f75d12018e91f8fe4

Observation a16053e1-5fbe-4422-b118-183795266389 · outbound

This paper cites an unresolved cited work.

JudgeBench: A Benchmark for Evaluating LLM-based Judges Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-18T01:24:02.132058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:24:02.052803Z digest=sha256:dac1cd44f8405e2fa1f96421fb364c0212cec875e97964e009763dd2772b615d

Observation af812f28-b742-4c25-96a3-9916772928c7 · outbound

This paper cites So, the final decision is Response 1 / Response 2 / Tie.

JudgeBench: A Benchmark for Evaluating LLM-based Judges So, the final decision is Response 1 / Response 2 / Tie

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T01:24:02.136211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:24:02.052803Z digest=sha256:e171585c966a30a94e79903db4f877d0ee1799561ede7f0024dddb1ae8087274

Pith citing papers

Observation 877b95d4-104b-4d9a-b190-241e22a2d80d · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 142

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:35:44.149506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:dfbc8c345608928210ca870c84348d2b2210e1ffb7cfe3346787b5722fda034a

Observation a44a7667-9683-416d-a720-4c3a3a793636 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 220

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:24:02.137278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:0ba406c800a3e8d85c645edb3ee28912423435d3e3d31919c104ad1bd0c8f97e

Observation 1127d743-065d-4dd2-a4b3-fe8c51514e52 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:24:02.137278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:4d688c8e04174692af19106c55e4a5e376062fdaaa223b134ff3971c691e0fcc

Observation ac8d15c4-c40b-430a-a911-560fb805d8a8 · inbound

Enhancing COBOL Code Explanations: A Multi-Agents Approach Using Large Language Models cites this paper.

Enhancing COBOL Code Explanations: A Multi-Agents Approach Using Large Language Models JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:40:03.461298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:40:03.461298Z digest=sha256:acb154b26f31460cd21d2a299418dacbdb45eb89f35dafbd2679963bd11d1481

Observation 68938448-2a4b-4ea9-98d3-3cad8e1b3909 · inbound

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards cites this paper.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.531419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.531419Z digest=sha256:99cb3714445abf863af6b5a3a58b2356f1ad3084c53b29ddd66728855402d56e

Observation 1ee9cd37-ef97-44e2-a5ff-b6d9c35525db · inbound

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution cites this paper.

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T22:54:25.184597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:54:25.184597Z digest=sha256:b12fd682f852411446b3d3af910c2947fdbd3110e2916295fef818696d858a52

Observation bbdbe3dd-7eae-40c2-b34e-81e813040f4a · inbound

NVIDIA Nemotron 3: Efficient and Open Intelligence cites this paper.

NVIDIA Nemotron 3: Efficient and Open Intelligence JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 184

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T01:40:42.750774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T01:40:42.190369Z digest=sha256:9dc4fb044237c77f5e3f3725d481142d55cda6be2adf99dfbb181b98cd5bc431

Observation cd097b4a-db0f-4555-af2d-6ea21860b988 · inbound

AIDG: A Formal Decomposition of Information Extraction and Containment Asymmetries in Multi-Turn LLM Dialogue cites this paper.

AIDG: A Formal Decomposition of Information Extraction and Containment Asymmetries in Multi-Turn LLM Dialogue JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T22:17:31.148047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:17:31.148047Z digest=sha256:3f2db70573dd4fea75f61136dfd0ea95fb901e4662aba8f0fe11675080c6a625

Observation b50371a8-365e-43e5-8085-354baf51de63 · inbound

MM-tau-p$^2$: Persona-Adaptive Prompting for Robust Multi-Modal Agent Evaluation in Dual-Control Settings cites this paper.

MM-tau-p$^2$: Persona-Adaptive Prompting for Robust Multi-Modal Agent Evaluation in Dual-Control Settings JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:24:02.137278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T13:43:52.671745Z digest=sha256:dc8cbc6fb5e89216a119036f2a76a44c5cb62fbaebea87b881570c447229765b

Observation 12cebc39-8826-4396-97af-2baf8d07d4d1 · inbound

Structured Multi-Criteria Evaluation of Large Language Models with Fuzzy Analytic Hierarchy Process and DualJudge cites this paper.

Structured Multi-Criteria Evaluation of Large Language Models with Fuzzy Analytic Hierarchy Process and DualJudge JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:24:02.137278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T17:12:07.852282Z digest=sha256:8b4cc65631a616931624f647a7d856c581ebde5494b841013688b1b8512b88bb

Observation 3ad28c37-11bd-4c01-8427-06b41a8221bf · inbound

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems cites this paper.

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:24:02.137278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T14:56:00.449776Z digest=sha256:d12d1366acf37fd4f2ad5596539b5688c52a228e00f0b46c4a90a397ed1f40ec

Observation 612a4f1e-701e-41a8-85f5-43f1c6a57e61 · inbound

Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering cites this paper.

Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:24:02.137278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T07:29:03.994957Z digest=sha256:1c3493514a6e7692c1c71771f2b138950414941a406d7e1db65c116160fe9622

Observation e24a3111-f640-4a69-a58e-70922bd160f1 · inbound

Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity cites this paper.

Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:24:02.137278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T11:45:54.181019Z digest=sha256:8c71d0c821c7f66d41c1f9fe3eb2e7933f25d49c0532815b72f079b2f6647f94

Observation 92b89e58-b68e-4f01-b646-8839c3ab19e7 · inbound

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines cites this paper.

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:24:02.137278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T08:14:18.535385Z digest=sha256:e5f356601b4a30c8a726863e6b61d5bacdfddbfa329a9ca10cddbccf505adf05

Observation 2c430462-4d8b-4fca-9ac5-7798a5fc2c2c · inbound

Green Shielding: A User-Centric Approach Towards Trustworthy AI cites this paper.

Green Shielding: A User-Centric Approach Towards Trustworthy AI JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:24:02.137278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T03:43:54.896449Z digest=sha256:c936041bb4ca761fa9ad32a27e8a367cb9bbdc7d9ca04f02db9fea9b992cf08d

Observation 182870d9-e3c5-4940-8fe9-9f5de1b3dbe1 · inbound

Training Computer Use Agents to Assess the Usability of Graphical User Interfaces cites this paper.

Training Computer Use Agents to Assess the Usability of Graphical User Interfaces JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:24:02.137278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T16:04:14.938406Z digest=sha256:889de1999064e946314d13bbbec21edfadadd1e11156dfd20c45dccd6bfcddce

Observation 7ca0f88c-7962-4e8e-a25a-3b8221eac6e6 · inbound

LATTICE: Evaluating Decision Support Utility of Crypto Agents cites this paper.

LATTICE: Evaluating Decision Support Utility of Crypto Agents JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:24:02.137278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T13:30:46.523784Z digest=sha256:450393e6198142c22714d9504024c8e314c3136bf6b1ae25d8be1e2d27d398b6

Observation 125ea4ac-d648-4461-8879-cc69a105ebb8 · inbound

LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding cites this paper.

LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:24:02.137278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T08:39:55.256518Z digest=sha256:3f88f482a8ec8f1732412795b9bcade07941198b4b044f9521f2df4556db23d1

Observation f24a1c1b-8d2a-4bfd-9a06-9682f53b81e8 · inbound

Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance cites this paper.

Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:24:02.137278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:50:50.037018Z digest=sha256:41323c4d072713a14025e7893bcda1bcb730c3d5b44a54061d2f49e3da86be43

Observation d612516c-8963-4972-aeea-dadb6b54358e · inbound

DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain cites this paper.

DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:24:02.137278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T02:40:27.234973Z digest=sha256:3f43d21e0f98e860a74eeb94f23a0329dbe2f03f4b95f29a7ab35fc5b342b994

Observation 9588355e-9fba-4ddb-9d11-f94ee8d924c6 · inbound

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge cites this paper.

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:24:02.137278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:19:00.238602Z digest=sha256:27a9fde699e7476b5b501f9ad26d92308207d88f55c5db351d0670373d5fe89b

Observation 828fc486-253a-4beb-9349-2a51aa29cb17 · inbound

RTLC -- Research, Teach-to-Learn, Critique: A three-stage prompting paradigm inspired by the Feynman Learning Technique that lifts LLM-as-judge accuracy on JudgeBench with no fine-tuning cites this paper.

RTLC -- Research, Teach-to-Learn, Critique: A three-stage prompting paradigm inspired by the Feynman Learning Technique that lifts LLM-as-judge accuracy on JudgeBench with no fine-tuning JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:24:02.137278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:37:09.912719Z digest=sha256:6e8ca7ce665aeee8383dff696ccca78749b7176230657cd636c2f65475c61f4e

Observation 6939b2af-23ce-4106-b771-cbac7f8f19a2 · inbound

Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering cites this paper.

Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:58:06.151710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T06:53:44.993529Z digest=sha256:425eeb9d39d43f91e5154e56b66dc9ad91cc56778ebb215aef91ad148dc2859d

Observation a8d17776-c9e9-4a32-821e-3c80a1799a81 · inbound

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts cites this paper.

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T19:02:34.054734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T19:01:14.340754Z digest=sha256:44d5ae56d945211e00d248b02765fc56abf77216ce8453224fbf1f7c78cca818

Observation 453009ee-8fc6-42b3-956d-c4809fabfdfb · inbound

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models cites this paper.

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:36:14.949514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T16:45:33.046568Z digest=sha256:1238448657c94527aa1b7aa8698959661bd98df2bf5c7eaff23d8d488c6292d8

Observation ef892dda-8de7-4890-a369-2d2ce812d574 · inbound

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks cites this paper.

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:36:27.442787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:46:24.554332Z digest=sha256:5adb57ae521120eb5dbbdae984d62ef9172740bf127634006dcf132690c9dada

Observation aa4d5695-7367-4495-9499-ae4c0f70d1b7 · inbound

RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning cites this paper.

RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T02:17:34.934879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:02:04.749003Z digest=sha256:a5b178a98f903581a6d517fc5e5bca567a7518eb0fcb409e3cea85d74d55e8d9

Observation 2074f077-bf95-4e0c-8087-3f56d7c4542b · inbound

Are LLMs Bad at Moral Reasoning? cites this paper.

Are LLMs Bad at Moral Reasoning? JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T13:18:12.657972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T08:20:24.251540Z digest=sha256:4e5499723fb995a76db241b27ce6612f5078a83e455edb791e8a3651b7dd416c

Observation 39937189-1cc9-47a8-b7b5-d222bdd21fd4 · inbound

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results cites this paper.

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 94

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T16:58:43.295766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:45:20.445703Z digest=sha256:714313f20a52332067bb77581e5a09d2dc0be4792eab62c7e182b25797c84c87

Observation fa124f31-66f9-4cc6-ac3d-d3e7e9e6d4b2 · inbound

Mind Companion: An Embodied Conversational Agent for Process-Based Psychotherapy cites this paper.

Mind Companion: An Embodied Conversational Agent for Process-Based Psychotherapy JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-03T22:59:03.273876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T23:07:48.991096Z digest=sha256:a4d048eb526c110e29859567ab7b997549e507c255231daba10018be095377dc

Observation 1e90c1f6-5d62-4bc6-a341-dd3dc89dab2f · inbound

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing cites this paper.

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-04T05:39:40.102134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T15:48:26.303462Z digest=sha256:928bc2140551793b56b402cc8dd95eaa7cd7f407e573ed2233e531d9e7650fe1

Observation 04a7c527-38c5-4e58-a62d-1ad662986a76 · inbound

Counsel: A Meta-Evaluation Dataset for Agentic Tasks cites this paper.

Counsel: A Meta-Evaluation Dataset for Agentic Tasks JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T06:49:38.225537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T14:07:59.446478Z digest=sha256:182bf70698fbdbe9171946d8bb7439193ba6bc5308dd8965d8f77b625fd062b8

Observation fcbb5378-2ff7-433e-9377-a60236cc6345 · inbound

Towards Spec Learning: Inference-Time Alignment from Preference Pairs cites this paper.

Towards Spec Learning: Inference-Time Alignment from Preference Pairs JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-07-04T11:39:47.085675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T07:49:36.816100Z digest=sha256:247aa2114928fbff578244b899c7be0a8c551a021655a69da2f0fe2ff20c1240

Observation 26100d5d-96b7-417b-8445-cdc38e93f0b2 · inbound

Towards Spec Learning: Inference-Time Alignment from Preference Pairs cites this paper.

Towards Spec Learning: Inference-Time Alignment from Preference Pairs JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-06-30T12:44:39.537486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T10:17:33.176525Z digest=sha256:45dce39a893a3d78559a5d914b2be3d29e958a9836ec8219a8fb1ebadcbe334a

Observation a43f1ec4-f362-47d9-a625-26365f72cbd2 · inbound

MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery cites this paper.

MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T17:09:58.528009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T23:57:26.000650Z digest=sha256:2df0e1cf0f67f98703dbd4101478b36cda541288e585dd6bd43495c331c025b3

Observation 3ccb3788-f29f-45a7-9e4d-5b475ef4c534 · inbound

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support cites this paper.

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-07-01T12:35:43.863245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T01:57:54.065453Z digest=sha256:75a60d98d2f51352decbd772bf97b844ec5df66f97c1109e0225c77ccb2e6185

Observation 6ac25137-7d17-4883-aef9-1070d6762439 · inbound

Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis cites this paper.

Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T07:46:14.714824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:46:14.714824Z digest=sha256:e2c0a2f328dc927e9daa246c4639ee05455d5d61526001a79002ede43b15d0ef

Observation 676058d8-0a36-4e19-bad9-c42b2862976d · inbound

TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning cites this paper.

TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-30T15:14:05.914715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T15:14:05.914715Z digest=sha256:b7d3852e815ce8421c77c926dd465ec0c7a7ef3add09fc1f268d264cade5639e

Observation 1abe627d-4e0d-4c7c-854b-c78201c607fd · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.143931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:26.143931Z digest=sha256:b66ce3a180adc359ca517fabbc31f746ac6b3a07cf33a6395074b368a6290cb1