Pith. sign in

Paper Citation Record · LEDGER

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2504.15253.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15253 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:56:01.605280Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:49:38.244210Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9718b30f-8129-4090-b209-9e0b43a91073 · inbound

Can You Trick the Grader? Adversarial Persuasion of LLM Judges cites this paper.

Can You Trick the Grader? Adversarial Persuasion of LLM Judges Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T21:56:01.605280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:56:01.605280Z digest=sha256:e7d187632b56ad6d8d3e0d47e1a68ba66272a1f95a4f89d9a9e9e155d6d9030d

Observation aa498442-b8c9-46d0-9379-96303932e058 · inbound

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models cites this paper.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:56.146103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:56.146103Z digest=sha256:6ef7ce5895a45cf3268fe2c7aef60dc88c641726b6d06aa27e03cfae5f844a44

Observation d3531d6c-a718-48ad-b4dd-4bfeb6ed5f0a · inbound

Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning cites this paper.

Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:47.902627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:39:47.902627Z digest=sha256:c5339cb224036e8de056cbaccb7f4353b6d1091881b4e0969cff86336d92a7e5

Observation 2052e7dc-8e34-4643-960f-36bb1c65a4ba · inbound

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization cites this paper.

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:56:24.588950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T12:53:45.767341Z digest=sha256:6f4a156ca60f8900ba878c45a0a9cf5882faa9d2ef6b7b58d8d721d3f2f2d65f

Observation 142f02d3-7d4c-4ff0-abbd-df389318e56e · inbound

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models cites this paper.

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:36:14.991198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T16:45:33.046568Z digest=sha256:494f597585ff8d9bb43c0b1b32fa77a9a24eb1c078c9f14daf4fa1be324d1508

Observation d2d5591c-75b5-4c20-a51b-0e8e55a72838 · inbound

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges cites this paper.

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:36:47.679091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T05:58:59.870335Z digest=sha256:828433e015e206b2f9c9ea5f715b5f0c980c316211ede3fd4d6fdac3200cc46a

Observation df095008-aa51-4e87-9f08-e3532340a4ae · inbound

Counsel: A Meta-Evaluation Dataset for Agentic Tasks cites this paper.

Counsel: A Meta-Evaluation Dataset for Agentic Tasks Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:49:38.246250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T14:07:59.446478Z digest=sha256:eb14bc7b24edd6ecb97f67fe7bde9945024437aa347fff93816e2ec51afa90bc

Observation d44d8c1f-a5a3-4d80-88a5-19b8ac73ceec · inbound

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification cites this paper.

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T02:43:21.225324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T02:43:21.225324Z digest=sha256:ede75bad09a90458b5e9f233e174dfdab74cf9400c8dc078951607954ad56cdb

Observation cf7f5c99-4b5e-4db2-9b3a-f92db4a927ac · inbound

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning cites this paper.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.038013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.038013Z digest=sha256:22a798f570833d4dfb10f14b9a9d5365537b1722a30dacf9046890ca7b910821