Pith. sign in

Paper Citation Record · LEDGER

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge

As of 17 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 1 inbound Pith citation observation for arXiv:2505.11875.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11875 v1

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:50:10.809590Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T06:51:56.556213Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T06:54:01.037769Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5ac5df32-1d18-42bc-a1be-30d764a83e2a · outbound

This paper cites Critique-out-Loud Reward Models.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Critique-out-Loud Reward Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.436494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.436494Z digest=sha256:ca5eb87c581e04a19db3c3ec5986c2a7012452d7e2af9cdea167dbbf83aa61e0

Observation b0790348-3993-483a-ae30-0118665643a4 · outbound

This paper cites Chain-of-Thought Reasoning In The Wild Is Not Always Faithful.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.444035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.444035Z digest=sha256:71b406f5cac01d40e01d5762fa2ab8c25ae2e583b65c0cfcef8916128b108f9f

Observation 651d2625-62df-4db1-a08b-cae1ffdd979b · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.449658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.449658Z digest=sha256:899387e51c139a0b63f049ee5ca14bc4f5ca658be0f09525e021495f278afcb3

Observation 88d6640e-8e54-4ec7-9bc1-7af1024a6ddc · outbound

This paper cites Rank analysis of incomplete block designs: I.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Rank analysis of incomplete block designs: I

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.455694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.455694Z digest=sha256:b48be84c665a98b260db2beea1bedde9eaf8d99980c885bd8b11c79bd1c7629e

Observation b8e11bc1-ab12-473a-9276-a6e1e919ceab · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.461007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.461007Z digest=sha256:0f5b8ab3086e055699c707f17ce08a9fdcbee8aa1d7d5d843521bf60ef336c06

Observation 14106471-aeec-4471-b4e2-9063e597965b · outbound

This paper cites Internlm2 technical report, 2024.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Internlm2 technical report, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.466439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.466439Z digest=sha256:d509a057962293a862eb74fec774765d146d460aa354c56939a787e60649fe03

Observation c1ac62e6-1f14-4b5f-8fa8-28400518939f · outbound

This paper cites RQ-RAG: Learning to Refine Queries for Retrieval Augmented Generation.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge RQ-RAG: Learning to Refine Queries for Retrieval Augmented Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.472032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.472032Z digest=sha256:95f929f46aef9d547ff4abcef4eef75721abdf5cc647489c28bbeb3410c65cc8

Observation 7f265e26-4bd5-407f-ac99-635c2e12cce4 · outbound

This paper cites CodeMonkeys: Scaling Test-Time Compute for Software Engineering.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge CodeMonkeys: Scaling Test-Time Compute for Software Engineering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.477308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.477308Z digest=sha256:ae2aba86ffa479e86e8a7785316eb7cf7d2c50980d7e097e2dbcfc0bd12949b3

Observation 59b2a579-93bb-4804-8422-89810d1e3aed · outbound

This paper cites Scaling laws for reward model overoptimization.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Scaling laws for reward model overoptimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.482771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.482771Z digest=sha256:d07c80d725e9058396850e6b87c8ddbd70bcd4dd956efb398b1544a312514b71

Observation 7d0b3ea7-ab84-403d-a478-a23e1727bb7c · outbound

This paper cites Gemini 2.0 flash thinking mode.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Gemini 2.0 flash thinking mode

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:50:12.190589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:50:10.487974Z digest=sha256:3dc3f03e0baddb8534a0a2ce01b45db20eca3fe0e7d4a59080d04e0576b2e00e

Observation 8a8d7b61-c5e5-4153-b04f-10555938505e · outbound

This paper cites The Llama 3 Herd of Models.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.493461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.493461Z digest=sha256:35a71b18a3fac9473053934093fd28a3cc94d473990b7c5cf5b058b11c18d4bb

Observation d9a0143e-6e3a-4eca-87a3-14a6d2fac6e0 · outbound

This paper cites A Survey on LLM-as-a-Judge.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge A Survey on LLM-as-a-Judge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.499201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.499201Z digest=sha256:b93a3174b05226c96ef17000ba003e5be3806599e70161896e940585ed63bdc9

Observation b173fb99-234d-4b52-ac4b-186ad1b9730b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.504807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.504807Z digest=sha256:796f015535e8c2cf0a68f4fa37a6d972887ff6ed21ffe90788a53a59957dc71a

Observation f1e6d08d-c4a8-4340-bfa6-6e477506c399 · outbound

This paper cites WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.509542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.509542Z digest=sha256:56d8ccd2e22522a341bbea581d9ed09575f4978ff3ae874b95741ac39f6d392a

Observation 3517ec8c-e637-46c0-8a66-19dd9032f937 · outbound

This paper cites Metrics for Explainable AI: Challenges and Prospects.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Metrics for Explainable AI: Challenges and Prospects

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.514533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.514533Z digest=sha256:9f15987b65ebb408aaa645576cc24d2d3bb2ac3877e32e3e370675862f5cb7ef

Observation cd9810c7-f516-48af-ba87-11efe248e784 · outbound

This paper cites Human Feedback is not Gold Standard.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Human Feedback is not Gold Standard

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.519624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.519624Z digest=sha256:f84db3d8ad237bf4082721a53c27544b4d77038131560b85671f74b3894efa8f

Observation 6f680da6-fc13-4fd1-afb2-0df8f6d4a6f3 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.524402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.524402Z digest=sha256:0c709492189f13c70645877c16855776410a55539770f25a694d8ec39533c961

Observation fc4e5462-60a0-42b2-a694-27af085da23e · outbound

This paper cites open-r1/openr1-math-220k.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge open-r1/openr1-math-220k

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:50:12.173715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:50:10.529827Z digest=sha256:b08bf9258cf88f4c416cbbafa001fcf6b2c57b3af22b6b42cdd5196aa2a35966

Observation 06262224-13c3-4f23-8a03-21b71a1463e5 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.534659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.534659Z digest=sha256:1bcd889d7640d56bd6058c369456f83ce8c627d7d4897d84aecaf0a78b0a7ee2

Observation c51ce50a-8322-42fd-a578-22b3ed8fc3f7 · outbound

This paper cites On scalable oversight with weak llms judging strong llms.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge On scalable oversight with weak llms judging strong llms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:50:12.156458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:50:10.539974Z digest=sha256:133d9e06b1c969606a04616036d8f26cd92f9ab5fbd30c512af45ff476e6b6dd

Observation ab2762c8-24bd-4359-9e61-9eb3ecee8bea · outbound

This paper cites Evaluating Robustness of Reward Models for Mathematical Reasoning.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Evaluating Robustness of Reward Models for Mathematical Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.544791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.544791Z digest=sha256:75386cdfd8be8ad3b7408b3fa558ed1a5d17e3ff5265b594ad25a663750555e5

Observation fcc46d3b-de50-4ce1-9122-010879d858a3 · outbound

This paper cites The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.549841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.549841Z digest=sha256:bc1735d14acf693cc90f43929057f668656626f0353f382641a81b2426ae05ed

Observation 5077b2c0-f17a-4ef2-aec8-ca995dad0fb5 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge RewardBench: Evaluating Reward Models for Language Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.554990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.554990Z digest=sha256:fa5080cc5dd00263496013f1a5473f643dcb7533490a956a1f51178d8c902b30

Observation d7f8acd1-148d-4dd3-86b6-a248966a084b · outbound

This paper cites From generation to judgment: Opportunities and challenges of llm-as-a-judge.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge From generation to judgment: Opportunities and challenges of llm-as-a-judge

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.559959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.559959Z digest=sha256:8908f1e12e1a2e7ccbcb6ecfe6ab3664ffa8e7e1d7e3916882fb4b9243c9699d

Observation b134077e-e4b5-48b2-90a9-8c8a634e2ff6 · outbound

This paper cites Generative Judge for Evaluating Alignment.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Generative Judge for Evaluating Alignment

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.564644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.564644Z digest=sha256:83efedf3c90be9ad723b044cf8ab747d770dc9c843cc5e9c78e78b2ce8deecbf

Observation 8d988751-2e87-461e-844d-882b2cf70639 · outbound

This paper cites Let's verify step by step.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Let's verify step by step

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.570180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.570180Z digest=sha256:85f616335f4043f1c8f36c8a1004f8011848b159862100375ee4d0bc432378f0

Observation 717df3ac-9d5c-4bfb-a2cd-60bbc516554e · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.575982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.575982Z digest=sha256:7082f4511af66714b71f046084d3d39b212f02c061d4b752221848a548b321b4

Observation fac3f3d2-a97d-49a0-947e-cbbe3f21fa50 · outbound

This paper cites Video-T1: Test-Time Scaling for Video Generation.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Video-T1: Test-Time Scaling for Video Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.580904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.580904Z digest=sha256:6fe901e5a7767165f74428c596a27983062ed3ce5e5d0f723405b1ed5ed662e3

Observation 74c64b38-a827-492c-a85e-aa5eb19e2e68 · outbound

This paper cites Learning Code Preference via Synthetic Evolution.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Learning Code Preference via Synthetic Evolution

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.586287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.586287Z digest=sha256:9a36096ba32ca8a93f54eb172e9a32e8dbaae8160c298daae7d58b39d4329f83

Observation 18096810-0622-4985-94e7-6b9917ed69c8 · outbound

This paper cites Inference-time scaling for generalist reward modeling.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Inference-time scaling for generalist reward modeling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.591957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.591957Z digest=sha256:3cdf946691b9c72ef1202e67e7e67bc4b952f2b0cda037428a38535b68bb8ced

Observation a53a6a3f-e3e1-4310-80bd-5f6d3cccfb02 · outbound

This paper cites Principal components analysis (pca).

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Principal components analysis (pca)

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:50:12.128796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:50:10.596839Z digest=sha256:29b518bba05a4e5ece37bb2b5b0fa152e143dcfc3032605f71e35a5b52c12cde

Observation 9c42e0fb-1f97-4bed-86b9-82d50605798e · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Self-refine: Iterative refinement with self-feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.601696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.601696Z digest=sha256:054347e5eec59506bd800544e5bf1e7f508674da93f969aae415ffdfab93a095

Observation ee70976a-dda7-4ce8-83d6-b3643658689c · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge LLM Critics Help Catch LLM Bugs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.606695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.606695Z digest=sha256:368c8f3b0cede07a8ee6a138f5e5634ecf116de42c7bbdbb3094d8240c3be2a8

Observation f4c9b583-8b7b-48a9-b3b6-37edf0b766e9 · outbound

This paper cites s1: Simple test-time scaling.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge s1: Simple test-time scaling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.612040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.612040Z digest=sha256:9116d15b866ed624c7545d06443b071d50130a865d2b21e0e10d59cfdd12811a

Observation 2534edf0-a286-426d-9b53-1d05f9dfa127 · outbound

This paper cites Learning to reason with llms.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Learning to reason with llms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:50:12.101504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:50:10.617303Z digest=sha256:e916a1ec4ee482115fb61d939d17a3a36aad1297b1e680ca8d11b8103602ffb6

Observation 74bf2cca-d301-45a0-91cb-cf6d7ee3cc84 · outbound

This paper cites Training language models to follow instructions with human feedback.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Training language models to follow instructions with human feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.622075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.622075Z digest=sha256:575a5359b742a03a46738c8e9498607f6896e94cff7fb279043df218ad97a387

Observation 9032f033-466d-4408-bd5f-3777967ffaeb · outbound

This paper cites OffsetBias: Leveraging Debiased Data for Tuning Evaluators.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.626884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.626884Z digest=sha256:87fcdc9b1b04b76f18a643b527ac77d72bc9917f4bf39c1f24566e65e65a650a

Observation 06f743cf-9983-4d01-b785-b09d267b93fe · outbound

This paper cites Codeforces cots.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Codeforces cots

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.632128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.632128Z digest=sha256:7f37a02a8073e49f9569197818ef5e2d04657dfaa4ca2369649367f54f82fcff

Observation 2dc70ec6-264d-45fc-83fe-70a0567d2b40 · outbound

This paper cites Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.636826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.636826Z digest=sha256:5831cf10fc14a920fb33b030e2b3ba77df804b81c1040f523b5626a659ae5873

Observation a3910a0e-0383-4bdb-9fc9-e02f56917c99 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.642154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.642154Z digest=sha256:7a6ca61c4b75cc81ba10357f8bb85046605dcba1a3f38550b3d8791d1219e032

Observation ab972d01-0af6-4a51-9f44-4d5429532e87 · outbound

This paper cites Proximal Policy Optimization Algorithms.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Proximal Policy Optimization Algorithms

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.647716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.647716Z digest=sha256:bd5769fb701dfebab645b87366fd8bf30956be41f056c39a4534df4bb1568750

Observation 10defcfb-39b0-4479-b126-91e9e4654881 · outbound

This paper cites Rethinking Reflection in Pre-Training.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Rethinking Reflection in Pre-Training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.656123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.656123Z digest=sha256:e51f76d90ec1593a6d3db860d5db966d6c8ec7bc69d743526cafa399d5b49a38

Observation 5265dbef-10ca-4f34-bacf-f96f10d32b45 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.661362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.661362Z digest=sha256:28a66485e07461ea98e31e55ffeb59f9b29c1933dac13ea14afa1b318dd9048a

Observation d7437445-4e5c-4204-b640-ffcd69daff0d · outbound

This paper cites Skywork critic model series.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Skywork critic model series

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:50:12.063946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:50:10.666183Z digest=sha256:9eb8dfac50d29e516cd3a4433c1d1f3db95b826d653dfb5a815aca540e18e260

Observation 0ccec354-002b-4a03-bac9-76eda568aa24 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.671068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.671068Z digest=sha256:1e6f00b7aaacd32b7bca2e92bf8abb9fd6e48ebdd349feab660224cedf3451ae

Observation 2eb1173f-22a9-49de-920e-8a77f649ecf5 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.676036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.676036Z digest=sha256:a943d177713af92169c15b1fad55e64aacf65654b850f9dc9463d63dd1ae2d73

Observation 7332358e-dafe-41d3-8598-01038e3d75a3 · outbound

This paper cites OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.680809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.680809Z digest=sha256:13d79d65f42a3484d9f8c9deb16c76e5fc5c9ea42b597cf6d0f79a78e5bc05f4

Observation 1d273eb4-3e0e-4360-aefa-53deaf8f8177 · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.685925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.685925Z digest=sha256:04744e0e6ad2e959421bb2fc1aa97ec12419a2d97afc9e35efd23ef43217ab86

Observation d1a924e3-b953-46bb-b261-45d2f70bc1f7 · outbound

This paper cites PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.690874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.690874Z digest=sha256:24d55b2208768a3500e7ff228938f10825840f9f1560cd8f7025ad6576e808bf

Observation 8f717129-d5ef-4671-9290-c40b8efb86ff · outbound

This paper cites HelpSteer2: Open-source dataset for training top-performing reward models.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge HelpSteer2: Open-source dataset for training top-performing reward models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.696490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.696490Z digest=sha256:f95c76a13eb760369d6b3920a06558ae40f1f2e9e46c5808bcf732049c7eab9c

Observation d2b6d259-7d69-40d3-8045-0e66ca820046 · outbound

This paper cites CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.701477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.701477Z digest=sha256:48b0a1591009b116c150e6aa4e1d038fabef8157bafeb9c1da55a745e3c52fd1

Observation 5b070bb3-c9d3-4433-99a9-60e9b48b9bce · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.706682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.706682Z digest=sha256:9d4be9b8ea8e2a5919dc304c903a2e952da01b73f0d8d32edc71c4aa8e5552ac

Observation 74837265-47e4-4644-a030-e78669015e47 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.711821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.711821Z digest=sha256:00fec004c8c885f5ef377983cead6ac2296ce21300a1d1f84c1ad19f6cc76751

Observation 2dfe73a8-18ed-42ac-ae77-95ff14473511 · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint, 2024.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.716804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.716804Z digest=sha256:301231e5483b7b9b3635903b3590eb665876cf8c3d5c427fa2b87ae15ec8d709

Observation 68bc86be-7a87-4501-9799-e0d7efb070ce · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.721529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.721529Z digest=sha256:206c3a7770aa725e5e8a00a76ff2644a8f61c73a0f606a021a2f8813ac61ebb7

Observation e8da5a41-b423-4431-af5e-73ffb094deec · outbound

This paper cites Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.727349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.727349Z digest=sha256:363b7578f449e91f1c5951adc12973bd8b378d5b5893dabd7ebd6419dc9e5368

Observation 27fb38c9-5622-4ad3-8c07-db05c783cf0c · outbound

This paper cites Qwen2.5 Technical Report.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Qwen2.5 Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.732498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.732498Z digest=sha256:5f5fb30655ac74dae3f39c1db8d30dd18051bcad07f35501d4888459a7b7fdc0

Observation da640ddb-bfa1-46a9-9be8-d2e574d7a34f · outbound

This paper cites Mastering complex control in moba games with deep reinforcement learning.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Mastering complex control in moba games with deep reinforcement learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:50:12.023617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:50:10.737309Z digest=sha256:398c51c9348f78bfffbb040d2704d4e7ec1801bf5d2c36c13b15acc26ea1d7b2

Observation 2420f695-7e96-4b88-8ad8-1164498e6519 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.742029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.742029Z digest=sha256:9e1e4b3d05d493d503132508838b7bfa00b91873df6f90e40711082bd14a5626

Observation f6e5db68-d593-4fe2-8b62-5192716d26df · outbound

This paper cites Improving Reward Models with Synthetic Critiques.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Improving Reward Models with Synthetic Critiques

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.747217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.747217Z digest=sha256:5166f3a16718c83b526e9fb8d921a81e21bcbb5aa687d407589e7c8e22a4162c

Observation fb263d60-1f64-4aa3-bf27-ec3c70ea846a · outbound

This paper cites Improve LLM-as-a-Judge Ability as a General Ability.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Improve LLM-as-a-Judge Ability as a General Ability

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.752460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.752460Z digest=sha256:f46c273591ae1183ba8aeb6669f77ce78fbb0fba458a4f8af77c76a75f7f0b33

Observation 13be2a22-8da4-46b1-9941-d6d1fb329489 · outbound

This paper cites Self-Generated Critiques Boost Reward Modeling for Language Models.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.757228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.757228Z digest=sha256:8875cf3fa531d7f6b69dcb2c68995f535afebc01b04895dc503209cea7f48ddc

Observation d38de029-615c-4ca0-aa7e-4a6ba0707f3f · outbound

This paper cites Z1: Efficient Test-time Scaling with Code.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Z1: Efficient Test-time Scaling with Code

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.762040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.762040Z digest=sha256:4233ed2fafb75c566dbb6161cf3d2864faf0ba1979639464c7bf1c1d40fe6e01

Observation 2b61f84a-557a-49f5-8e22-56a820d7d41e · outbound

This paper cites LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.767033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.767033Z digest=sha256:c27a672380b295f61de2d511ad25bf022ba92361c79686b629263fc70623a3b5

Observation cf0ccd04-6bfa-4465-af41-e84551740a41 · outbound

This paper cites Openprm: Building open-domain process-based reward models with preference trees.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Openprm: Building open-domain process-based reward models with preference trees

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:50:12.006003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:50:10.772283Z digest=sha256:2b1b8c4b3a901cfa03934b1a0993fb51b6668f03b315fe2c2726ae7aaed0fbfe

Observation 99e9c94b-caf6-494d-960e-65dde388b2a1 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.777366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.777366Z digest=sha256:4fe68fd3187c1a1d6c27951d3b067661f32ffdd3c922597b2413e37bf5b7f394

Observation 75580b05-339e-4015-b754-d08d2082f886 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.782262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.782262Z digest=sha256:9a85b8b1feb1bfdfdf59eed1eeac0d7acb64a064fa47a2c4d46e203fdb9273cd

Observation 3142b804-82d7-4317-a93c-a1dbceeb2f23 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.787220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.787220Z digest=sha256:4a73e32326e01f09276bff18216e00d6d4ca2d6038c263dd756347e3f8ed0631

Observation a68f265e-105c-4f61-8486-e1a9b497d9b1 · outbound

This paper cites write newline.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge write newline

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.792368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.792368Z digest=sha256:ac10e051ec8045e632f41f28934afaa9336d4bbef22a43135933e0a41c839927

Observation e2fcb828-fd43-413a-97e8-d86becababed · outbound

This paper cites @esa (Ref.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge @esa (Ref

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.798582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.798582Z digest=sha256:477b45c83e780e56e82020c1d65672289c2647e1f088731f380983c497be0234

Observation 789001e0-dfb7-45ef-828a-a5dc19afdac5 · outbound

This paper cites an unresolved cited work.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.804046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.804046Z digest=sha256:297337e2f5392c48d43aeed5a68ca642a66183bfcdc89894880c8a8d166be01e

Observation 2482c9fe-d73e-4b1d-bbef-857ee8f47697 · outbound

This paper cites wait" step across four benchmark tasks. Because the proportion of responses altered after adding each.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge wait" step across four benchmark tasks. Because the proportion of responses altered after adding each

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.809590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.809590Z digest=sha256:10f8f3d2dc5cdb7e9ffb279491dac3a77b4e7d3fb864c85c81a1a9b319c2da39

Pith citing papers

Observation 800902ac-7871-4cda-8b64-b03a5c706977 · inbound

The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering cites this paper.

The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:54:01.039517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T06:51:56.556213Z digest=sha256:66fb7790cc9550467cc33c949bbde426bde56f43ff5d203accf36cb5ac851036