Pith. sign in

Paper Citation Record · LEDGER

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

As of 8 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 7 inbound Pith citation observations for arXiv:2505.16265.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16265 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:56.857674Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:27:05.232691Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.265666Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c58bb891-0435-43ff-9642-7742c8b15d5b · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:52.682509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:52.682509Z digest=sha256:6c6edc7611bf5998a322454416d20693185f23591c75907894cbfabe4fdce010

Observation 59bfa91b-2909-4427-9ccd-c67d9ec13674 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:52.737506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:52.737506Z digest=sha256:7fea8ec8ad659f510822d44bb3848afb9a1cf3293a35c154d54ade5ea12fdf75

Observation f07ae36f-c76b-4d7c-8c1f-f222743a2dc6 · outbound

This paper cites Self-instruct: Aligning language models with self-generated in- structions.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Self-instruct: Aligning language models with self-generated in- structions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:52.808601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:52.808601Z digest=sha256:9aa60bf92fe9c7058912b54a6576b4499faf2ab0916b4f78eace5b6856881ff8

Observation 43a330cc-1fed-4d3c-9827-4ac2971ef4f5 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Direct preference optimization: Your language model is secretly a reward model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:52.870219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:52.870219Z digest=sha256:ddd9691a1600f512d697f8099cc1f4b1d8b31b9f9d3a1416cc55b21e9787ee36

Observation d72f5654-f8a9-48f7-8ff6-9ad313727da2 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Solving math word problems with process- and outcome-based feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:52.975342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:52.975342Z digest=sha256:4de9e2c4c6b454f9c0acbe6834b8404aa727b7faf5e3ee92bc80528eb70a7b01

Observation 923b156a-bd74-4bff-985b-e3b4802d60f3 · outbound

This paper cites Let’s verify step by step.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Let’s verify step by step

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.051311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.051311Z digest=sha256:cb8774a9bd56f6af4117ff719456ab189b37450378f3bd277a16c5c26a11d490

Observation a518297f-2337-45c0-949c-ee65ff909fff · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.131349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.131349Z digest=sha256:d0afc644bfdf96002a163718a2a2f5b225ee47fc52dd54c440f04936b9716699

Observation d19f54ef-db2a-4b5b-867e-f353c9648dd4 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.173407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.173407Z digest=sha256:38637c02cc28e53bdb71e148235d34e9e58d0fd70c2ae18958e7ee0951296a89

Observation 95b87412-9d1f-44a2-84a9-81f256e99745 · outbound

This paper cites Beavertails: Towards improved safety alignment of LLM via a human-preference dataset.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Beavertails: Towards improved safety alignment of LLM via a human-preference dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.615816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:53.261459Z digest=sha256:629eecd3d7694803c20d967f9cf9c9bee7ba6f664f6831fe2a476543c40f92f8

Observation b8da442f-c411-4031-8c19-ccdcebc0bbe0 · outbound

This paper cites Safe rlhf: Safe reinforcement learning from human feedback.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Safe rlhf: Safe reinforcement learning from human feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.325000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.325000Z digest=sha256:c71fbe16a7abc25eb276ee6a69b4db46611cde33638ade55371fba4c1091132f

Observation 1be1e9ba-9aab-4af3-9a66-735e7c64dbff · outbound

This paper cites Rule Based Rewards for Language Model Safety.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Rule Based Rewards for Language Model Safety

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.390130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.390130Z digest=sha256:6baa7d6aebc382bed2320cf48ea6079c4bf4aa6d60a637af46a1f01afb327761

Observation 6a542a66-92af-4f80-9787-089b613587b3 · outbound

This paper cites Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.452403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.452403Z digest=sha256:49f54502ad545743c0dd2387e168de9d32eed4eae70c4764617c25a472c160f0

Observation 9714b577-6cec-468a-b8b7-8700d92a58cb · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul F.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul F

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.370846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:53.539620Z digest=sha256:9870da3ed8c6e403da678de805c47d83f236c1311ff66dc3be521b091b453d32

Observation f9169d0a-e4b5-4e96-b42f-62d01dd8571c · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.635786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.635786Z digest=sha256:887fc9948ab1a87b9e84341d730c806d031c0b57c789c082e08d3561cf8e8a9c

Observation 2e3a2de9-5eb9-4149-98b7-fb67fb6bd51f · outbound

This paper cites Scaling laws for reward model overoptimization.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Scaling laws for reward model overoptimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.734461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.734461Z digest=sha256:359746031baf62be0073eea90b8b88e2804927328cf95672b6f86b991973e3f9

Observation 9cb827c2-9926-4f6e-8717-aabe5f2532a0 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.852749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.852749Z digest=sha256:bb922667f659bb51d81b4f936e9102e99895c1f652953f393d6f6c9459fbab81

Observation 31781822-1fb8-43b9-b342-af1f9554021d · outbound

This paper cites RRM: Robust Reward Model Training Mitigates Reward Hacking.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.954441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.954441Z digest=sha256:67b88346b81090f4735085daabda359acb109ab6d4d99a1e5ad665c9a2733163

Observation aa3293c5-7248-418f-b6d0-259228f7611e · outbound

This paper cites Improving Reward Models with Synthetic Critiques.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Improving Reward Models with Synthetic Critiques

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.077941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.077941Z digest=sha256:ae4a417196967cdc97d352c912c907bf588ee6d7b00f9324a0d7b7de15f6f115

Observation 73873e99-f590-43c3-8f88-36992110828b · outbound

This paper cites Critique-out-Loud Reward Models.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Critique-out-Loud Reward Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.211281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.211281Z digest=sha256:3c3dd3e973932445125a9b5879928f41251eaa74c2f045906520d4c8f2994dbc

Observation 02278e09-b24d-478b-8dae-36186b32ead0 · outbound

This paper cites Generative Reward Models.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Generative Reward Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.317164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.317164Z digest=sha256:0d03ee986994812229513373f4880e3eeaa4876ecec799fc73c5dedff10d7c7c

Observation faf6a9ce-0c08-4f47-8de0-a77a1dec0b40 · outbound

This paper cites Self-Generated Critiques Boost Reward Modeling for Language Models.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.445894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.445894Z digest=sha256:b330ebb37fdcfcdeb0e8128f5dfcbffc5a380363c205b1144eeec96e402e476b

Observation ab3d36a6-1f52-434c-a252-04f94a95dbe2 · outbound

This paper cites Generative verifiers: Reward modeling as next-token prediction.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Generative verifiers: Reward modeling as next-token prediction

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.140897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:54.584180Z digest=sha256:4da6eb94a9fe95ac4d581708b25956c54b46dc36fd1d7ba445f4213f2df3a359

Observation ca8cef68-5cf3-4bfc-9169-7caea7e48c55 · outbound

This paper cites Learning to reason with llms.OpenAI Blog, 2024.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Learning to reason with llms.OpenAI Blog, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.673142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.673142Z digest=sha256:0d156c0e640db880f9a3557e5c7fd4bac109396f51d8c50e0025a4fa2269c7c9

Observation c2ca967b-49ae-483d-a002-177acfefe3f3 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.768046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.768046Z digest=sha256:a26e671c0da3d2dcea6ff0b1bacd6b466bd3effe75daf93d9a18ddb3c0ccbcc4

Observation d056e02f-8513-4fec-955c-7b8b41f8f23e · outbound

This paper cites s1: Simple test-time scaling.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models s1: Simple test-time scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.919059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.919059Z digest=sha256:fa383050c06fa5a5bd048c7af9f7bb07ef672c5d23716ec56af1d892a1abe835

Observation d21f3082-9dce-4316-a9f9-2f06c95f2fcf · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models LLaMA: Open and Efficient Foundation Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.000866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.000866Z digest=sha256:1e7a9f2a8759ba611fcaa8aac7e13813f92944cc6b9a05cbe278e2a729da296f

Observation b3c836f0-fc4e-4263-b699-052ca9d012b6 · outbound

This paper cites The Llama 3 Herd of Models.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models The Llama 3 Herd of Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.146505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.146505Z digest=sha256:6fac7c0586dff2956f0047d7f8c85ba19444cbd7ce430f92273714523e38d33b

Observation d6f6d54b-6ec2-4a63-9475-00a3d0475250 · outbound

This paper cites RM-bench: Benchmarking reward models of language models with subtlety and style.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models RM-bench: Benchmarking reward models of language models with subtlety and style

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.283280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.283280Z digest=sha256:a1ae6b24df22585bd9ba2360506559114e647f506368786ffd2fbd9cae8efcec

Observation 00e257ad-035a-4abd-8868-118ac98e178b · outbound

This paper cites Rank analysis of incomplete block designs: I.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Rank analysis of incomplete block designs: I

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.398260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.398260Z digest=sha256:b7251a8d78cb1b98d701e24988db01a3900833d7ef8c55a24dcac7a1c832c20a

Observation 1f30d9f2-2268-4864-8233-10613ff33b72 · outbound

This paper cites GPT-4 Technical Report.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models GPT-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.488038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.488038Z digest=sha256:a9a9271ab20d3e94dbb96e84b03147a9896f97cd96e721747bbfffee7e57fd73

Observation 82111220-12b3-46b0-8fc2-b0f3719e97c0 · outbound

This paper cites Qwen2.5 Technical Report.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Qwen2.5 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.548298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.548298Z digest=sha256:0bf5999de2575418f04c73df1ebbb1ff901e46695688a135273f0beaafe1955e

Observation 315b6cf5-c3f8-4782-a3d8-6976fb258f4a · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.607193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.607193Z digest=sha256:60cd381ae9e59df8df65afa2f151e6188ca05ab14880a6afd18da4dc150aed02

Observation 405f8c7c-4d30-4e61-a0bd-a7d3597ff81a · outbound

This paper cites xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.688065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.688065Z digest=sha256:274ab43ce474821a736bb1619ab2f5907d61313bbb0d6712b631ce39223f3841

Observation ea497a77-b5f1-46f0-a748-54aa1e3ede75 · outbound

This paper cites Chain of thought prompting elicits reasoning in large language models.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Chain of thought prompting elicits reasoning in large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.870900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:55.756393Z digest=sha256:a0b39832c0a917962c194f7eb85f268a3921977026db4a065bbd7ea4b3dd367b

Observation 4755a606-83be-4f69-839a-7730e7adbc10 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.832125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.832125Z digest=sha256:0be7f982a2c2903f6fe125c76c47540739db9dd71ca9cb8df38606762d9b14d7

Observation 95e5ed7f-22b3-4919-bbdc-19ac14b174d8 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Ad- vances in neural information processing systems, 36:11809–11822, 2023.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Tree of thoughts: Deliberate problem solving with large language models.Ad- vances in neural information processing systems, 36:11809–11822, 2023

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.913418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.913418Z digest=sha256:c3ad345ff55e9a9e136629a7f36735bbc3b1a69eef5668ac3ca4b94579689771

Observation f8eedf22-3a99-45bf-8183-efc9e4022c63 · outbound

This paper cites Introducing openai o3 and o4-mini.OpenAI Blog, 2025.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Introducing openai o3 and o4-mini.OpenAI Blog, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.678509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:55.990614Z digest=sha256:7a9cca88efe6a4517e98dd1c6b20bb067f84e200af1ed186dbdf427462bf0a67

Observation 91ec9f65-4299-4d48-995d-13cef4560e11 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.047246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.047246Z digest=sha256:feda1c11dd7dd9ee82becccfde022aa44f365c24cd0b77ae422db72f60681bbd

Observation fef25670-483a-49d5-8d98-06f2cf72795a · outbound

This paper cites Grok 3 beta — the age of reasoning agents.xAI Blog, 2025.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Grok 3 beta — the age of reasoning agents.xAI Blog, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.514382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:56.122255Z digest=sha256:4affdfcb22b81f595bfac72b98ff70b3140fd0a28c878891648df1ff435db1cf

Observation 1f226cff-6000-4620-a171-35e03b575a8b · outbound

This paper cites Helpsteer2-preference: Complementing ratings with prefer- ences.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Helpsteer2-preference: Complementing ratings with prefer- ences

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.371923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:56.198046Z digest=sha256:a8f2c28939f1ea157856d4f57af5ad6f4469ef68554b9ed111620a53ac6e1b38

Observation 491e81c0-9044-40e1-ae49-84b255cc435f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Proximal Policy Optimization Algorithms

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.275062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.275062Z digest=sha256:01f14789a7a714422cb71423cf415d7cb04ff180fe6b37ab055f2fdd2d0b668e

Observation f9970fc8-6749-4400-8c26-6721eeff6864 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.314481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.314481Z digest=sha256:5bfca452c652723c2bde54b48de36a8ec1898a6d28e05c42c67461c6e9ea8c4c

Observation 2c5bc38d-4ee2-48c6-b992-dabbb58dff82 · outbound

This paper cites HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.386550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.386550Z digest=sha256:09c3d589e77e949dd89650ffb52fee777a879dc1cbfda7737b390c102a3d28f8

Observation c77ff6d2-21a3-4b1c-b875-fabd16585d9a · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models RewardBench: Evaluating Reward Models for Language Modeling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.474821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.474821Z digest=sha256:0dac332d4cb4c12a2c706f6f90799aaf200239bde276c344b56af9c064c47f04

Observation ff40040a-91c2-4b89-acc5-7b8adba6ef66 · outbound

This paper cites Length-controlled alpacaeval: A simple debiasing of automatic evaluators.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Length-controlled alpacaeval: A simple debiasing of automatic evaluators

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.236959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:09:56.552982Z digest=sha256:0068308854cd0c2030f88e51b4b929af3a518507bcc697e082853b74ec7ce94c

Observation f22dc59e-be31-420c-83ff-cf2f8ae7d59d · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.640514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.640514Z digest=sha256:2d42cc463749a263bfb052258b2b0cf58b383b99035c3d3fbb5def3488bb8b91

Observation 855b4e30-ca65-46b9-8145-2edb3eeeaab0 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models HybridFlow: A Flexible and Efficient RLHF Framework

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.727030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.727030Z digest=sha256:7373bb130657c8ea7e3cf40bc665fca76199e373a950d6f694f1bf84ebf0c35c

Observation 3f293609-f5a0-4863-ace7-1a3fb58bb197 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Adam: A Method for Stochastic Optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.803561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.803561Z digest=sha256:9f9dde16a7488133f1de71dc73014e3cf6a394c67833e0e22d3a4946dbbb1c80

Observation 6a916405-9325-47ba-adc5-217306390a62 · outbound

This paper cites Decoupled Weight Decay Regularization.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Decoupled Weight Decay Regularization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.857674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.857674Z digest=sha256:05b616830aa4469d81a2185ab75021f904c62edf535dd6ddb8fe95171b9f6712

Pith citing papers

Observation 09ba2b4f-5426-456b-8e8d-e3b90054b011 · inbound

VRPRM: Process Reward Modeling via Visual Reasoning cites this paper.

VRPRM: Process Reward Modeling via Visual Reasoning Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:21:30.982861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T12:20:17.430881Z digest=sha256:1a5cdf57f49688a65160d2ef576ffb7b59c5747681a721b820d1a52e137c146a

Observation d56e645e-ecad-49e2-beb1-9b4a70e64d7c · inbound

VRPRM: Process Reward Modeling via Visual Reasoning cites this paper.

VRPRM: Process Reward Modeling via Visual Reasoning Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T04:27:05.232691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:27:05.232691Z digest=sha256:739bbb207586bd25a9f96fa93335bfedd8236ba6cdf46f927eb52aaaa14c08aa

Observation a4e76bdc-41c2-4d56-a2ec-563800eacb62 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 266

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T19:21:48.839803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:b2327cc60fb1d556330a8b44f7477c4ed0795849b5312db30e5306af4bed2832

Observation d543c8c8-7392-4f7c-a4c7-18eb6997aa46 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 196

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:05:31.521389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:6330321f59ea82726b585257322d9b4b7056e33e9b1f39e320af6ac8ba6a5ec4

Observation 6a7b75e1-8549-4c77-a49b-f46d21f74ffc · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:06:27.574471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T08:00:33.307429Z digest=sha256:7e6a5981451d23703a3f1351b3d1924d93c8cb1cc6d80d3b2972af959ea4c98f

Observation 7c2feb22-2a50-40a3-a289-dbdae529330b · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:14:06.015503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T09:11:02.183133Z digest=sha256:f3c89fae95048296048059cb511954a13033eba3da249e3fbb6b4b8a4ba4f365

Observation 1c86b97d-afc4-47f6-b8d9-8d25caf5fcec · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 153

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.267006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:36d25410f03b20685ded6a31663d48514f6f3bf03d0e81053cd8489867758eca