Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:56.857674Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 7 inbound Pith citation observations for arXiv:2505.16265.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:56.857674Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:27:05.232691Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T20:56:13.265666Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c58bb891-0435-43ff-9642-7742c8b15d5b · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59bfa91b-2909-4427-9ccd-c67d9ec13674 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f07ae36f-c76b-4d7c-8c1f-f222743a2dc6 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Self-instruct: Aligning language models with self-generated in- structions
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43a330cc-1fed-4d3c-9827-4ac2971ef4f5 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Direct preference optimization: Your language model is secretly a reward model
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d72f5654-f8a9-48f7-8ff6-9ad313727da2 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Solving math word problems with process- and outcome-based feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 923b156a-bd74-4bff-985b-e3b4802d60f3 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Let’s verify step by step
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a518297f-2337-45c0-949c-ee65ff909fff · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d19f54ef-db2a-4b5b-867e-f353c9648dd4 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Teaching Large Language Models to Reason with Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95b87412-9d1f-44a2-84a9-81f256e99745 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Beavertails: Towards improved safety alignment of LLM via a human-preference dataset
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b8da442f-c411-4031-8c19-ccdcebc0bbe0 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Safe rlhf: Safe reinforcement learning from human feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1be1e9ba-9aab-4af3-9a66-735e7c64dbff · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Rule Based Rewards for Language Model Safety
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a542a66-92af-4f80-9787-089b613587b3 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9714b577-6cec-468a-b8b7-8700d92a58cb · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul F
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f9169d0a-e4b5-4e96-b42f-62d01dd8571c · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e3a2de9-5eb9-4149-98b7-fb67fb6bd51f · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Scaling laws for reward model overoptimization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cb827c2-9926-4f6e-8717-aabe5f2532a0 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Direct Language Model Alignment from Online AI Feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31781822-1fb8-43b9-b342-af1f9554021d · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models RRM: Robust Reward Model Training Mitigates Reward Hacking
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa3293c5-7248-418f-b6d0-259228f7611e · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Improving Reward Models with Synthetic Critiques
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73873e99-f590-43c3-8f88-36992110828b · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Critique-out-Loud Reward Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02278e09-b24d-478b-8dae-36186b32ead0 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Generative Reward Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faf6a9ce-0c08-4f47-8de0-a77a1dec0b40 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab3d36a6-1f52-434c-a252-04f94a95dbe2 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Generative verifiers: Reward modeling as next-token prediction
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ca8cef68-5cf3-4bfc-9169-7caea7e48c55 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Learning to reason with llms.OpenAI Blog, 2024
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2ca967b-49ae-483d-a002-177acfefe3f3 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d056e02f-8513-4fec-955c-7b8b41f8f23e · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models s1: Simple test-time scaling
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d21f3082-9dce-4316-a9f9-2f06c95f2fcf · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models LLaMA: Open and Efficient Foundation Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c836f0-fc4e-4263-b699-052ca9d012b6 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models The Llama 3 Herd of Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6f6d54b-6ec2-4a63-9475-00a3d0475250 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models RM-bench: Benchmarking reward models of language models with subtlety and style
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00e257ad-035a-4abd-8868-118ac98e178b · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Rank analysis of incomplete block designs: I
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f30d9f2-2268-4864-8233-10613ff33b72 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models GPT-4 Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82111220-12b3-46b0-8fc2-b0f3719e97c0 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Qwen2.5 Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 315b6cf5-c3f8-4782-a3d8-6976fb258f4a · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 405f8c7c-4d30-4e61-a0bd-a7d3597ff81a · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea497a77-b5f1-46f0-a748-54aa1e3ede75 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Chain of thought prompting elicits reasoning in large language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4755a606-83be-4f69-839a-7730e7adbc10 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95e5ed7f-22b3-4919-bbdc-19ac14b174d8 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Tree of thoughts: Deliberate problem solving with large language models.Ad- vances in neural information processing systems, 36:11809–11822, 2023
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8eedf22-3a99-45bf-8183-efc9e4022c63 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Introducing openai o3 and o4-mini.OpenAI Blog, 2025
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 91ec9f65-4299-4d48-995d-13cef4560e11 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fef25670-483a-49d5-8d98-06f2cf72795a · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Grok 3 beta — the age of reasoning agents.xAI Blog, 2025
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1f226cff-6000-4620-a171-35e03b575a8b · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Helpsteer2-preference: Complementing ratings with prefer- ences
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 491e81c0-9044-40e1-ae49-84b255cc435f · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Proximal Policy Optimization Algorithms
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9970fc8-6749-4400-8c26-6721eeff6864 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c5bc38d-4ee2-48c6-b992-dabbb58dff82 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c77ff6d2-21a3-4b1c-b875-fabd16585d9a · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models RewardBench: Evaluating Reward Models for Language Modeling
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff40040a-91c2-4b89-acc5-7b8adba6ef66 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Length-controlled alpacaeval: A simple debiasing of automatic evaluators
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f22dc59e-be31-420c-83ff-cf2f8ae7d59d · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 855b4e30-ca65-46b9-8145-2edb3eeeaab0 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models HybridFlow: A Flexible and Efficient RLHF Framework
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f293609-f5a0-4863-ace7-1a3fb58bb197 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Adam: A Method for Stochastic Optimization
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a916405-9325-47ba-adc5-217306390a62 · outbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Decoupled Weight Decay Regularization
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09ba2b4f-5426-456b-8e8d-e3b90054b011 · inbound
VRPRM: Process Reward Modeling via Visual Reasoning Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d56e645e-ecad-49e2-beb1-9b4a70e64d7c · inbound
VRPRM: Process Reward Modeling via Visual Reasoning Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4e76bdc-41c2-4d56-a2ec-563800eacb62 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
Reference 266
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d543c8c8-7392-4f7c-a4c7-18eb6997aa46 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
Reference 196
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6a7b75e1-8549-4c77-a49b-f46d21f74ffc · inbound
Leveraging Verifier-Based Reinforcement Learning in Image Editing Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7c2feb22-2a50-40a3-a289-dbdae529330b · inbound
Leveraging Verifier-Based Reinforcement Learning in Image Editing Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1c86b97d-afc4-47f6-b8d9-8d25caf5fcec · inbound
Trust Region On-Policy Distillation Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
Reference 153
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.