Pith. sign in

Paper Citation Record · LEDGER

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

As of 22 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 10 inbound Pith citation observations for arXiv:2504.15253.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15253 v2

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:33:53.636506Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:49:05.104352Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:49:38.244210Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved67
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 451c1f70-5522-4b0f-9ee9-fc2aaf2dd917 · outbound

This paper cites write newline.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.300764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.300764Z digest=sha256:bda4160e3629baa9e797f66ca66c3cbc6613f178dad48063556ca9e347eafe45

Observation 1e018f6e-37b3-4206-a847-e0be8b76d140 · outbound

This paper cites Program Synthesis with Large Language Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Program Synthesis with Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.306561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.306561Z digest=sha256:2208499d4543da753eda9dc04d7f6e5791eb25bc0240a209b9e8e289b20b15db

Observation d3fd40bc-5e7f-479e-adc7-90325d5e8077 · outbound

This paper cites Graph of thoughts: Solving elaborate problems with large language models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Graph of thoughts: Solving elaborate problems with large language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:33:54.855145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:33:53.311394Z digest=sha256:75a85a7861aeef403a68202a9e43e74b252c1329b3e89b79a7a2b288bed12fc0

Observation 8d9a5256-1a74-4b01-89b0-92406c8ab486 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.315974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.315974Z digest=sha256:bb803b6e22bfe5ce77a7ce9684f0094fe96e300f3d66356355ba4d0e50e903ae

Observation 6734b43b-77e0-4d0f-87b0-b27b138c6bd4 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Evaluating Large Language Models Trained on Code

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.320892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.320892Z digest=sha256:c0022ee22d5e2abb1aaf69ce7cf93514e5ea36767a5d634f777ed1c8a1b63d73

Observation 0e9e4198-24f9-4b71-b9e6-4a3fb9f7094b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.325919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.325919Z digest=sha256:09b57a4dd0b3c400e29221945927f9d10a93d573de81198a6b96ed6cb43330bc

Observation 5382001d-e052-4af1-a101-b649e9e57838 · outbound

This paper cites Process Supervision-Guided Policy Optimization for Code Generation.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Process Supervision-Guided Policy Optimization for Code Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.330823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.330823Z digest=sha256:43e143afe35d53f1ac1b9d3f876c98b77ddb0e15756ddbc0f908f70b8a448248

Observation 46968a37-e714-4b95-81dd-993715acf818 · outbound

This paper cites GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.336123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.336123Z digest=sha256:ca17c7669b6e12260ffba9da8ba9bb2349cfb8917575a2bbf9a086c1b709674d

Observation 8e0849b1-8c21-4331-84d1-6e3a8d375610 · outbound

This paper cites The Llama 3 Herd of Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.341100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.341100Z digest=sha256:3e9fee11058108cc4ce2ffd76f8a1be51faa7a201cfdd61e638f4dbd57db5b8b

Observation e6daa37d-52c6-474b-acd6-ec7fa12234bb · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.345382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.345382Z digest=sha256:cd44a34b44c3878411d5c80485ec99dddb17e51b864c560715e76deea77c169f

Observation 85ac88c3-18a2-4a08-8e08-2f353d058c93 · outbound

This paper cites X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:33:54.839906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:33:53.349928Z digest=sha256:fe40c250f3aaf79a00e5926ed700691ef45ce1ee187d079c4dc987e5544e02df

Observation 7675ace2-b13c-44df-9582-dab965298b83 · outbound

This paper cites Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.354266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.354266Z digest=sha256:7aa2684e12ccc1753b79c71b18f9e6035591abd71a89024fdbc659aa0b251de1

Observation 245b1fac-f3ff-4cb2-bdb8-45ead85e4fab · outbound

This paper cites How to Evaluate Reward Models for RLHF.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators How to Evaluate Reward Models for RLHF

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.358755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.358755Z digest=sha256:821308fb836d4aa984946371f28110c1a4772646929303f250e82425b317ff0f

Observation 718aeae7-a6f7-48db-bb6b-75809036aa84 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.363247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.363247Z digest=sha256:8e5205a4221595755710dea9b8b42cd1d2b93585c574ccf0bc9147cb6a414f8b

Observation 94484cd0-cd2e-4ee2-96f2-51cbea4c2f9d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Measuring Mathematical Problem Solving With the MATH Dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.367809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.367809Z digest=sha256:c7ffca9b7178f6e478023a08891b6a8d678ea36c4946e18698cc326d0484d6e1

Observation 50da763f-3f6e-4d99-ae73-32b655d2f706 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Training Compute-Optimal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.372411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.372411Z digest=sha256:86db77ebc581413c4f5f786a8e85c973d67ee2f5eeeecc712d26ac8615ba89d2

Observation b2fbcddc-be2e-4b00-9041-1ceb52025062 · outbound

This paper cites Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.377306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.377306Z digest=sha256:badaf7425ea4e97a66d96cc51b0697067856c2a9df05969a729817b5c72e2d35

Observation e5274ca1-0998-49b3-bad0-771f4fb2b6e2 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Large Language Models Cannot Self-Correct Reasoning Yet

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.381904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.381904Z digest=sha256:910195db33a8bc2922655e2b269e24e07caf1962a325ccdd12c5c0f306171341

Observation 16cfd505-a884-41f9-bfed-d064cdbc8927 · outbound

This paper cites OpenAI o1 System Card.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators OpenAI o1 System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.386368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.386368Z digest=sha256:b0333860d6f8987478c5edded580659ec66b9bb6bfd51492755ddaa208fa80bc

Observation 98edb3db-bb0c-4ce8-8863-150c35e6c302 · outbound

This paper cites A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.390529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.390529Z digest=sha256:ba198aa0ee5035b5f5ac2e25fe0b106f1311667106eb4629dcba1aa813976481

Observation cd580d81-d1da-463a-8886-9c18146aec37 · outbound

This paper cites Scaling Laws for Neural Language Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Scaling Laws for Neural Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.395112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.395112Z digest=sha256:cd00bd9183b99b1f4fce9c0503e72b03d13ba15cea441476f7b972a943fe4d7d

Observation 741aaaec-75c1-418c-ac2c-524de9216fc7 · outbound

This paper cites X., Li, M., Qin, C., Wang, P., Savarese, S., et al.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators X., Li, M., Qin, C., Wang, P., Savarese, S., et al

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.399161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.399161Z digest=sha256:dbf7486ce42e4ec82667425ff605885e592b2ac33938259b2546e0bc6d7cb031

Observation 478c77ed-6022-478c-98c7-a2628de02832 · outbound

This paper cites Prometheus: Inducing fine-grained evaluation capability in language models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Prometheus: Inducing fine-grained evaluation capability in language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.403828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.403828Z digest=sha256:e4585191975bab7113311294d16e900fa7d7803225a7fd956873304e7ff64677

Observation 3598cd01-4af1-4f62-99a6-b6773f0a2feb · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.408060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.408060Z digest=sha256:b50ea214b7d6bb11c0eb6a6b8c2860593bb2c711032d96f43336e6058afce743

Observation f0f36a76-4d25-4b65-a4a3-121f6a7bd1fe · outbound

This paper cites S., Reid, M., Matsuo, Y., and Iwasawa, Y.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators S., Reid, M., Matsuo, Y., and Iwasawa, Y

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.412740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.412740Z digest=sha256:d4d0f3fb5855de1da580a5a24a2e618ceeee59a713c9c638fc50422f2f1857f5

Observation 06c7b997-4740-48d4-a1d8-939997b73ae0 · outbound

This paper cites H., Gonzalez, J.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators H., Gonzalez, J

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.416910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.416910Z digest=sha256:aae17b9bbf65347b28d9ec3d62f6e39914724940412ffa835eff049ad5a8eb85

Observation a8d4e9df-5494-4c13-b0ba-be468bf2a113 · outbound

This paper cites Math-Verify: Math Verification Library , 2025.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Math-Verify: Math Verification Library , 2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:33:54.798859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:33:53.421111Z digest=sha256:1a7d53c310dacf6d3af23819bce4af0ffeac02ecec955cdd9df062a9254d2553

Observation aaa7a588-d331-4ce9-9af7-3d02a7a0495b · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators RewardBench: Evaluating Reward Models for Language Modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.425258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.425258Z digest=sha256:9baba9c3acb095ba77d1edf4a1017e1231f2db1cd2e512bc2a75247c44d2c936

Observation 24a687f5-59d2-46b2-a294-3d73cfbc7814 · outbound

This paper cites CriticEval: Evaluating Large Language Model as Critic.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators CriticEval: Evaluating Large Language Model as Critic

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.429585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.429585Z digest=sha256:7b2505131cac3ff51303934e106d5db2c06810f24755364a27f24892e1888ec2

Observation 0c69488b-b291-4319-ab87-4ff79f3be329 · outbound

This paper cites Generative Judge for Evaluating Alignment.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Generative Judge for Evaluating Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.433729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.433729Z digest=sha256:feebecec4b2e27ed42da8933b5006050184893d1a6e3f56381acc90d6b54a339

Observation f58363e7-5549-46c9-83a7-01634021b3ce · outbound

This paper cites Let's Verify Step by Step.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Let's Verify Step by Step

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.438021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.438021Z digest=sha256:b0ec919e74b9e9075f128b4a9f73ef3d4ed1f43f3016a9613fa3ab30be4ab03c

Observation e31cffcc-7df2-4230-b851-9b85a6631491 · outbound

This paper cites CriticBench: Benchmarking LLMs for Critique-Correct Reasoning.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.442371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.442371Z digest=sha256:edecb484325c22ff04d816c9a1a69fbb50d62a350ab2a20675ae598b32b38579

Observation 55611e9e-14d5-4da4-ba8f-e0885cb0dfe2 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.448043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.448043Z digest=sha256:7a631a1fa9b78d425232338842cd36ff40970a186cb8deedb772873120aebaa5

Observation 496ae6bc-ffce-44ee-a05d-039d8c7124ef · outbound

This paper cites S., Wang, Y., and Zhang, L.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators S., Wang, Y., and Zhang, L

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.452426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.452426Z digest=sha256:ff57257d78da66e38e317c78cfb26880fefd5e0ee0153c892e56d304cca3d2a9

Observation 643549e0-a18c-401e-8fba-9d944bfb4a08 · outbound

This paper cites Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.456653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.456653Z digest=sha256:d561e677395d3b65827eb255464959990b44668a10fc1dc98095f0440a92ade7

Observation 78b95e4f-913e-434f-af6e-f7fc95ba8fb5 · outbound

This paper cites RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.460702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.460702Z digest=sha256:62d132e8dafba6508204be3b4ef14a6df0fa17f27e0fbcb0cd8d202f1777459c

Observation 2567b4ed-d673-475c-82dc-496b3fe088e1 · outbound

This paper cites PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.465051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.465051Z digest=sha256:7f3e2be940f713cc191bee9f355d3f32db05e40ec74c86e3db36f6d36830e828

Observation 0c122cdd-c3a1-4556-844c-e53606f84df6 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.469382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.469382Z digest=sha256:d657f789db282d505a88098c8d4d407b61537b921b28f3f38f428bedb0442367

Observation 3d1c0f5d-5c49-4186-89e1-780c76352aa5 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Self-refine: Iterative refinement with self-feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.473889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.473889Z digest=sha256:3737ba9e803e8aac50cce4146184ff0634345b30898e9cdb36284bfd4251d0cd

Observation 977c2e83-8fa3-429a-a89e-285fad4c092e · outbound

This paper cites CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.478015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.478015Z digest=sha256:afd6cb7afd349cdf477f104d1d49884baa9b2c6b141bb8c4309dd1466eddf517

Observation 7191b780-1b64-44d8-9b56-ba9b3a3e6bdd · outbound

This paper cites Show Your Work: Scratchpads for Intermediate Computation with Language Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Show Your Work: Scratchpads for Intermediate Computation with Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.482766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.482766Z digest=sha256:375b03a11b89f7b3bc482098c8ae6bc10fe3a95a3bda8c951ca88772406314f7

Observation 82918c6a-4387-4887-a60e-c49d60dbb4db · outbound

This paper cites Training language models to follow instructions with human feedback.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Training language models to follow instructions with human feedback

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.487338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.487338Z digest=sha256:e7cd9f0432d9f57ced27a27179e7e786a50c7d94bddf1219fe0fb97d35a8bdb8

Observation c8014eb0-9e0e-4c9d-b102-143e4b4e20af · outbound

This paper cites OffsetBias: Leveraging Debiased Data for Tuning Evaluators.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.491812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.491812Z digest=sha256:2972f6a79e572c80124d456d5a0a8fbecc201d2a9b0194e0758ac863159292b1

Observation 7d945d56-b8ce-42cc-bd2d-e87388d2bd76 · outbound

This paper cites S., Franklin, M., Vidgen, B., Singh, A., Kiela, D., and Mehri, S.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators S., Franklin, M., Vidgen, B., Singh, A., Kiela, D., and Mehri, S

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.496376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.496376Z digest=sha256:5f33a4ab793c039ff0cd0a69264e260bbbeeb6b9feaad1494181a59feb4c34d4

Observation de9f9b56-54bd-4a22-83e2-e57183e165ac · outbound

This paper cites Self-critiquing models for assisting human evaluators.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Self-critiquing models for assisting human evaluators

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.500649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.500649Z digest=sha256:d49a8c184f5d5b1936325d84dd5d46f4e00bc05ecd3f84c7d882e2729b1dfbfd

Observation c7149a45-f235-4072-ac3d-1408ab792ee6 · outbound

This paper cites B., Balakrishnan, S., Bradley, J., Parekh, A., Ramch, K., Wainwright, M.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators B., Balakrishnan, S., Bradley, J., Parekh, A., Ramch, K., Wainwright, M

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.505284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.505284Z digest=sha256:a0fff8389b80741aed81062b3e0ae0783fb3172d32bfe76b8f8c0d6b1d6900cd

Observation 84c7b553-33c8-4038-b7ca-83bf8d5f3c8c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.509484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.509484Z digest=sha256:d705b8585ddd0dda84d31ce56c8e849f295b23bea81e4dc8dac687e8aee269e2

Observation 60870029-d586-48a4-9a55-6e9654e0b148 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Reflexion: Language agents with verbal reinforcement learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.513992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.513992Z digest=sha256:89307f2a81761ba831d46300c7768bed8500522b228aa1d6da89732fbc63090e

Observation a1f1f110-c69c-4b07-956a-f2a978274924 · outbound

This paper cites Y., Zeng, L., and Liu, Y.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Y., Zeng, L., and Liu, Y

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:33:54.741694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:33:53.518677Z digest=sha256:3f6e56496304432392d280195847a979cf41a8bdebc75c279584fb9a67a28b82

Observation 6859ef88-2e65-4791-8061-39c3d1502403 · outbound

This paper cites The art of llm refinement: Ask, refine, and trust.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators The art of llm refinement: Ask, refine, and trust

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:33:54.728380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:33:53.522865Z digest=sha256:d0d35b9dd7ad063f0dd7da33dc4e377f233fe8a8f74be43bcc6e5e94f35ef2e7

Observation 3dece122-5448-4b6d-8dbc-e0a8d63d92ba · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.527127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.527127Z digest=sha256:f9a7416c114ede243fad615832e45700752c584e0e416019a8a641738c97e4d0

Observation 536088e7-b448-45e6-9de9-e38a6752366d · outbound

This paper cites GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.531866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.531866Z digest=sha256:ca36fb2e61f38db33bf5542e609c707fb381c5e04291fe89eac4707d698d4a69

Observation e18be681-439a-4905-a4ba-f8122d176cc5 · outbound

This paper cites JudgeBench: A Benchmark for Evaluating LLM-based Judges.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.536375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.536375Z digest=sha256:a17d8b4306b4e91ffc479a6cab655d272be637ef40ca9b4cf37598878ad1901c

Observation 954c0c8f-311b-4be5-9824-a8c4696ce09c · outbound

This paper cites Can Large Language Models Really Improve by Self-critiquing Their Own Plans?.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.540998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.540998Z digest=sha256:f85523d2c59693355f6e4990b469528751955dca07b18573fe76ba9ae6eff855

Observation 4396bc91-83a8-44e2-88cc-fb980e50cc79 · outbound

This paper cites Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.545529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.545529Z digest=sha256:08398ab43515c65f5fbec3d5bd04e6e47ee611b4956dc761dfba5e5549695d2e

Observation 5e6632f4-8e24-4e34-b0ad-ea2e400f7683 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.550130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.550130Z digest=sha256:7ef3329ecb79e36f1ac65fe2e7a776dcea85f64ec78303c746c7f8e51d86b035

Observation c5ef18c6-baba-4b8d-9333-4ecf6eb80d1b · outbound

This paper cites Direct Judgement Preference Optimization.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Direct Judgement Preference Optimization

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.555545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.555545Z digest=sha256:e16c4d7e78ef58c6ef9e7d779476756301722d111997784b5761cdf2f89f5e56

Observation 701a70b4-f8a6-47a9-bc3b-267b2ab298cb · outbound

This paper cites Self-Taught Evaluators.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Self-Taught Evaluators

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.560109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.560109Z digest=sha256:4233037031ff85780fecdaf5eed3826647801f3ac5974ee6124efb6ab49770ba

Observation 44ad93b7-01c2-4723-b5da-506d0a383821 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.564625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.564625Z digest=sha256:8c0a55c39e4528bdd521753cfbcdb427226a1008dcdf01b6af276f788babba52

Observation f43c3f82-067e-49ec-8574-7cbffc4cc8e9 · outbound

This paper cites Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:33:54.714084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:33:53.569347Z digest=sha256:c326245ae81aea341f9b496b66d14641dfb3763de9527dfdcd9eba42ec28a7a1

Observation 2ffd2a87-60c8-4488-b6b3-6768ed2f5116 · outbound

This paper cites V., Zhou, D., et al.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators V., Zhou, D., et al

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.573865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.573865Z digest=sha256:003e97426e4d64332ffda000458db68a8b661d32b79553b783735baa587bc76c

Observation 7690cba8-e0a8-4132-9453-0385b6bbbd3d · outbound

This paper cites Qwen2.5 Technical Report.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Qwen2.5 Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.578079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.578079Z digest=sha256:bf84c7a3668a6bbbb6fee5e994d90d179960d60c6e721bc343f02663fdccb138

Observation 0d165790-bc5c-4919-acc8-6e354c6a6d38 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Tree of thoughts: Deliberate problem solving with large language models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.582331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.582331Z digest=sha256:b154c0cfcb10bbb6becb4b3962b165ce59f7208e08a73d07f1390432d3ce3fac

Observation 5b3dfb1a-361f-4133-8fb0-96b2811fb289 · outbound

This paper cites Beyond Scalar Reward Model: Learning Generative Judge from Preference Data.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Beyond Scalar Reward Model: Learning Generative Judge from Preference Data

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.586506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.586506Z digest=sha256:4d3e338a19b410a6bd70b86a367e86b86d7a2575ddee60e076130912d8623b95

Observation 53dcbc1c-be50-4ba6-951b-5dc1a141dbc5 · outbound

This paper cites Y., Cho, K., Li, X., Sukhbaatar, S., Xu, J., and Weston, J.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Y., Cho, K., Li, X., Sukhbaatar, S., Xu, J., and Weston, J

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:33:54.681184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T11:33:53.591291Z digest=sha256:540d796436292105a4c5f46aa8d01c9555e8be39f7dbfc3d40c9f9f90e02205d

Observation eb76a554-c718-4ac2-857d-9205a4eaa57e · outbound

This paper cites Evaluating Large Language Models at Evaluating Instruction Following.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Evaluating Large Language Models at Evaluating Instruction Following

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.595367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.595367Z digest=sha256:506edf4026b82695970b6056903f1c600893ef5b904b149b76cc0a16c9a05f16

Observation 95280075-68d4-4a70-8c31-ed4e79f19255 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.600017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.600017Z digest=sha256:2d6b70c484c4ca17881cb4b7e6cb5113b8b6596ed0497841cd46fdb451f66f02

Observation 54fb66db-55dc-43d6-a631-c2d3b0562de9 · outbound

This paper cites Small Language Models Need Strong Verifiers to Self-Correct Reasoning.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Small Language Models Need Strong Verifiers to Self-Correct Reasoning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.604669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.604669Z digest=sha256:a6e829903930023cbf6dd75af4a11d64991429501fe2ae14502a5b54a75b6686

Observation 90a44f28-9f90-41be-ae81-addd9d3b8ecc · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.613435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.613435Z digest=sha256:405e703466a8ae6e561cffbebcb1879eb80e76d193d63a436c0e46f01d9ef459

Observation a4613a37-9cde-4a00-8718-f664a494ef58 · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.617714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.617714Z digest=sha256:fbb90f801a3a2428650c43d3ff76565ca6268565d84c6a597b645136664db153

Observation a5994195-4f4c-4b6c-966e-692c32e0dd98 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.622552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.622552Z digest=sha256:dbc34a11888d53fbe5ffaa1dc1462fd9313d1ec016cc4a56bdb6256d7cd9f61f

Observation c567d069-316d-455a-9c12-4d8d5de7d696 · outbound

This paper cites RMB: Comprehensively Benchmarking Reward Models in LLM Alignment.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.626881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.626881Z digest=sha256:3e935ca5bd129139b4cb9f1872cad44b887ac1764a402c1ed727e40dd738918e

Observation 72a44a73-3fdc-4843-bbf8-c12095820ac7 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators Instruction-Following Evaluation for Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.631721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.631721Z digest=sha256:e4343e8f6127c7f1e3cf5a51b9584832debefef140c5d88d6a8d265c88113f5d

Observation 3c1b5ab9-6c50-4923-8389-96ab9666bb70 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:53.636506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:33:53.636506Z digest=sha256:ce597be2b09a3ae2a745801d16332c07c23c270ef0db730dae457c00c070100e

Pith citing papers

Observation 9718b30f-8129-4090-b209-9e0b43a91073 · inbound

Can You Trick the Grader? Adversarial Persuasion of LLM Judges cites this paper.

Can You Trick the Grader? Adversarial Persuasion of LLM Judges Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T21:56:01.605280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:56:01.605280Z digest=sha256:eaf22862765db12f89935999058e4abd2d1c1ab3702b8b8619f80f30708fdde9

Observation aa498442-b8c9-46d0-9379-96303932e058 · inbound

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models cites this paper.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:56.146103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:56.146103Z digest=sha256:314f8ce68c127abccf6efc6fc8db2e256c987000a09f41f76f628892185456e8

Observation d3531d6c-a718-48ad-b4dd-4bfeb6ed5f0a · inbound

Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning cites this paper.

Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:47.902627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:39:47.902627Z digest=sha256:b0d1f167cfd99a3d776d2059502b8989b8d80e046864e6930642217146a3c036

Observation 2052e7dc-8e34-4643-960f-36bb1c65a4ba · inbound

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization cites this paper.

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:56:24.588950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-18T12:53:45.767341Z digest=sha256:0928f136fe924074d5841d8ecd3ba65c1ac45f11129e82a917f7716aa48c29ea

Observation 142f02d3-7d4c-4ff0-abbd-df389318e56e · inbound

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models cites this paper.

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:36:14.991198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T16:45:33.046568Z digest=sha256:da0d05e04824f73b1c011f85d335a68ce0101e9e898cfc835395129cb6312f21

Observation d2d5591c-75b5-4c20-a51b-0e8e55a72838 · inbound

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges cites this paper.

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:36:47.679091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T05:58:59.870335Z digest=sha256:49b26f701c028a34254734b5e34a806ecd03773e7774104ea1778002a9bba1a4

Observation df095008-aa51-4e87-9f08-e3532340a4ae · inbound

Counsel: A Meta-Evaluation Dataset for Agentic Tasks cites this paper.

Counsel: A Meta-Evaluation Dataset for Agentic Tasks Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:49:38.246250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T14:07:59.446478Z digest=sha256:097d03cee3ba2e461fdea696ecb588c83d080ccdf2fb7d86593a4402ed18a300

Observation d44d8c1f-a5a3-4d80-88a5-19b8ac73ceec · inbound

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification cites this paper.

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T02:43:21.225324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T02:43:21.225324Z digest=sha256:2ab849e473c915d55d35520170fa7871aaa33d97b4af1d579b9fbb9b3734974d

Observation cf7f5c99-4b5e-4db2-9b3a-f92db4a927ac · inbound

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning cites this paper.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.038013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.038013Z digest=sha256:5aeaa2249d61021dc9f84657b4906666529a187a18e5175845090838c4c3a3c1

Observation 82ca2942-d1db-4f9d-aa2d-2b2159e19bfe · inbound

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility cites this paper.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 151

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:05.104352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:05.104352Z digest=sha256:4d4710c0128dbfc16946686f979ef429e9c0c83a3da47ff3e3cffbeb855cdcb6