Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:50:10.809590Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 1 inbound Pith citation observation for arXiv:2505.11875.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:50:10.809590Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-21T06:51:56.556213Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T06:54:01.037769Z
72 of 72 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5ac5df32-1d18-42bc-a1be-30d764a83e2a · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Critique-out-Loud Reward Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0790348-3993-483a-ae30-0118665643a4 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 651d2625-62df-4db1-a08b-cae1ffdd979b · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88d6640e-8e54-4ec7-9bc1-7af1024a6ddc · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Rank analysis of incomplete block designs: I
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8e11bc1-ab12-473a-9276-a6e1e919ceab · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14106471-aeec-4471-b4e2-9063e597965b · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Internlm2 technical report, 2024
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1ac62e6-1f14-4b5f-8fa8-28400518939f · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge RQ-RAG: Learning to Refine Queries for Retrieval Augmented Generation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f265e26-4bd5-407f-ac99-635c2e12cce4 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge CodeMonkeys: Scaling Test-Time Compute for Software Engineering
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59b2a579-93bb-4804-8422-89810d1e3aed · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Scaling laws for reward model overoptimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d0b3ea7-ab84-403d-a478-a23e1727bb7c · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Gemini 2.0 flash thinking mode
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8a8d7b61-c5e5-4153-b04f-10555938505e · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge The Llama 3 Herd of Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9a0143e-6e3a-4eca-87a3-14a6d2fac6e0 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge A Survey on LLM-as-a-Judge
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b173fb99-234d-4b52-ac4b-186ad1b9730b · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1e6d08d-c4a8-4340-bfa6-6e477506c399 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3517ec8c-e637-46c0-8a66-19dd9032f937 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Metrics for Explainable AI: Challenges and Prospects
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd9810c7-f516-48af-ba87-11efe248e784 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Human Feedback is not Gold Standard
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f680da6-fc13-4fd1-afb2-0df8f6d4a6f3 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc4e5462-60a0-42b2-a694-27af085da23e · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge open-r1/openr1-math-220k
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 06262224-13c3-4f23-8a03-21b71a1463e5 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c51ce50a-8322-42fd-a578-22b3ed8fc3f7 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge On scalable oversight with weak llms judging strong llms
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ab2762c8-24bd-4359-9e61-9eb3ecee8bea · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Evaluating Robustness of Reward Models for Mathematical Reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcc46d3b-de50-4ce1-9122-010879d858a3 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5077b2c0-f17a-4ef2-aec8-ca995dad0fb5 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge RewardBench: Evaluating Reward Models for Language Modeling
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7f8acd1-148d-4dd3-86b6-a248966a084b · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge From generation to judgment: Opportunities and challenges of llm-as-a-judge
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b134077e-e4b5-48b2-90a9-8c8a634e2ff6 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Generative Judge for Evaluating Alignment
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d988751-2e87-461e-844d-882b2cf70639 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Let's verify step by step
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 717df3ac-9d5c-4bfb-a2cd-60bbc516554e · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fac3f3d2-a97d-49a0-947e-cbbe3f21fa50 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Video-T1: Test-Time Scaling for Video Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74c64b38-a827-492c-a85e-aa5eb19e2e68 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Learning Code Preference via Synthetic Evolution
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18096810-0622-4985-94e7-6b9917ed69c8 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Inference-time scaling for generalist reward modeling
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a53a6a3f-e3e1-4310-80bd-5f6d3cccfb02 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Principal components analysis (pca)
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9c42e0fb-1f97-4bed-86b9-82d50605798e · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Self-refine: Iterative refinement with self-feedback
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee70976a-dda7-4ce8-83d6-b3643658689c · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge LLM Critics Help Catch LLM Bugs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4c9b583-8b7b-48a9-b3b6-37edf0b766e9 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge s1: Simple test-time scaling
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2534edf0-a286-426d-9b53-1d05f9dfa127 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Learning to reason with llms
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 74bf2cca-d301-45a0-91cb-cf6d7ee3cc84 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Training language models to follow instructions with human feedback
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9032f033-466d-4408-bd5f-3777967ffaeb · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge OffsetBias: Leveraging Debiased Data for Tuning Evaluators
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f743cf-9983-4d01-b785-b09d267b93fe · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Codeforces cots
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dc70ec6-264d-45fc-83fe-70a0567d2b40 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3910a0e-0383-4bdb-9fc9-e02f56917c99 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab972d01-0af6-4a51-9f44-4d5429532e87 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Proximal Policy Optimization Algorithms
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10defcfb-39b0-4479-b126-91e9e4654881 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Rethinking Reflection in Pre-Training
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5265dbef-10ca-4f34-bacf-f96f10d32b45 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7437445-4e5c-4204-b640-ffcd69daff0d · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Skywork critic model series
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0ccec354-002b-4a03-bac9-76eda568aa24 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb1173f-22a9-49de-920e-8a77f649ecf5 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7332358e-dafe-41d3-8598-01038e3d75a3 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d273eb4-3e0e-4360-aefa-53deaf8f8177 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1a924e3-b953-46bb-b261-45d2f70bc1f7 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f717129-d5ef-4671-9290-c40b8efb86ff · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge HelpSteer2: Open-source dataset for training top-performing reward models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2b6d259-7d69-40d3-8045-0e66ca820046 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b070bb3-c9d3-4433-99a9-60e9b48b9bce · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74837265-47e4-4644-a030-e78669015e47 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dfe73a8-18ed-42ac-ae77-95ff14473511 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint, 2024
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68bc86be-7a87-4501-9799-e0d7efb070ce · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8da5a41-b423-4431-af5e-73ffb094deec · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27fb38c9-5622-4ad3-8c07-db05c783cf0c · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Qwen2.5 Technical Report
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da640ddb-bfa1-46a9-9be8-d2e574d7a34f · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Mastering complex control in moba games with deep reinforcement learning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2420f695-7e96-4b88-8ad8-1164498e6519 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6e5db68-d593-4fe2-8b62-5192716d26df · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Improving Reward Models with Synthetic Critiques
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb263d60-1f64-4aa3-bf27-ec3c70ea846a · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Improve LLM-as-a-Judge Ability as a General Ability
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13be2a22-8da4-46b1-9941-d6d1fb329489 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d38de029-615c-4ca0-aa7e-4a6ba0707f3f · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Z1: Efficient Test-time Scaling with Code
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b61f84a-557a-49f5-8e22-56a820d7d41e · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf0ccd04-6bfa-4465-af41-e84551740a41 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Openprm: Building open-domain process-based reward models with preference trees
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 99e9c94b-caf6-494d-960e-65dde388b2a1 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75580b05-339e-4015-b754-d08d2082f886 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3142b804-82d7-4317-a93c-a1dbceeb2f23 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a68f265e-105c-4f61-8486-e1a9b497d9b1 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge write newline
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2fcb828-fd43-413a-97e8-d86becababed · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge @esa (Ref
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 789001e0-dfb7-45ef-828a-a5dc19afdac5 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Unresolved cited work
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2482c9fe-d73e-4b1d-bbef-857ee8f47697 · outbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge wait" step across four benchmark tasks. Because the proportion of responses altered after adding each
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 800902ac-7871-4cda-8b64-b03a5c706977 · inbound
The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.