Pith. sign in

Paper Citation Record · LEDGER

Reward Reasoning Model

As of 8 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 15 inbound Pith citation observations for arXiv:2505.14674.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14674 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:53.874909Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:15:37.528799Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:49:38.201535Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2b5dc76e-1feb-4983-8214-fa8e19d1cf8a · outbound

This paper cites GPT-4 Technical Report.

Reward Reasoning Model GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.218433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.218433Z digest=sha256:7506e4734cd03c20fe7e4595251daaa8a74edebacf05094870ce8140be57c7cf

Observation ef46c4eb-d007-43dd-bcac-af4fe3fbe2f0 · outbound

This paper cites Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models.

Reward Reasoning Model Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.264050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.264050Z digest=sha256:e522177b2100099789a9533888d810098c80af869a973ff00029dcca5fe47d45

Observation 10d6e181-a23c-4326-a1ea-b681e34819c8 · outbound

This paper cites Atla Selene Mini: A General Purpose Evaluation Model.

Reward Reasoning Model Atla Selene Mini: A General Purpose Evaluation Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.298845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.298845Z digest=sha256:c4b5c0a153f875da90881ea1a7a0700cd9cdae866756c7e10ee5d6fb85388a07

Observation 324dd137-dbb8-4b26-ad1d-bfb5723149e7 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.Claude-3 Model Card, 1:1, 2024.

Reward Reasoning Model The claude 3 model family: Opus, sonnet, haiku.Claude-3 Model Card, 1:1, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.335857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.335857Z digest=sha256:b51e2bdc5be71d52833baf810499cd33b8850e00d9bdc5e6068d8b64c0892752

Observation abb5fcfe-f89d-4712-9a0a-c907f47bcfec · outbound

This paper cites an unresolved cited work.

Reward Reasoning Model Unresolved cited work

Reference 5

Resolution
verified exact
doi, observed 2026-08-07T15:34:54.013847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:50.375780Z digest=sha256:57be531a02809fdc55cadcf96a644637343f32005faababb28a88e4c4dc835a9

Observation 47b45be1-fd2f-4c65-b61d-eb409a76aef7 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Reward Reasoning Model Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.406624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.406624Z digest=sha256:96cb0594040d271ff90d63bf96d92b5dca868034ef17c71ed02e9d0f79b527db

Observation f6cb77be-ec9f-4514-991a-2a89a49521dc · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Reward Reasoning Model Constitutional AI: Harmlessness from AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.446245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.446245Z digest=sha256:41d6577aca78f9dcb9c861ef056e428d138a9dbe85bdbc67b49eec1c814c177e

Observation c9e244ad-e186-419b-a6b6-3427b84bc06e · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Reward Reasoning Model Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.491218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.491218Z digest=sha256:a3d4a2f38fc4630e5c085d4cb9ab93b5ddede9ab8419b2ef63d165e1626d2df0

Observation 7d2dc476-2b95-4d48-8f9e-b5fce7fb796a · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

Reward Reasoning Model Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.535585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.535585Z digest=sha256:01ff2d54139a42315af300893295b30d2f30524645afd2dfa5cfdef348ea0691

Observation b2be30d7-3095-4200-b422-7a1795082215 · outbound

This paper cites Sparks of artificial general intelligence: Early experiments with GPT-4, 2023.

Reward Reasoning Model Sparks of artificial general intelligence: Early experiments with GPT-4, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.592474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.592474Z digest=sha256:7414239b6c1c249a65258282955336d0e7ffa5ac6f2b257946cae7f9afa2836e

Observation f46c7d44-57de-479f-95fe-2c2c9e4dd482 · outbound

This paper cites CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution.

Reward Reasoning Model CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.638608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.638608Z digest=sha256:ec6ecb6d600eea1c55c0d181fa24f7d8487b70b058ac3c87c4cb1ab1c0fe0924

Observation ecaf28c3-59e1-4e5f-9512-4f3d20541476 · outbound

This paper cites Judgelrm: Large reasoning models as a judge.arXiv preprint arXiv:2504.00050, 2025.

Reward Reasoning Model Judgelrm: Large reasoning models as a judge.arXiv preprint arXiv:2504.00050, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.684754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.684754Z digest=sha256:78801a10e883885bd877d91775ec97c151b8b1b2caac7281677b60de0dd88ce5

Observation a6981fcf-1d7b-42ff-b33b-3da231f42faf · outbound

This paper cites Seal: Steerable reasoning calibration of large language models for free.

Reward Reasoning Model Seal: Steerable reasoning calibration of large language models for free

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.725027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.725027Z digest=sha256:6397d534c1d0147679164e8b9bcd006181bd3918eab0c30d4bf672d4008dd59c

Observation 89dbad2b-12a3-4097-97aa-2fe0d4550f18 · outbound

This paper cites Rm-r1: Reward modeling as reasoning.arXiv preprint arXiv:2505.02387, 2025.

Reward Reasoning Model Rm-r1: Reward modeling as reasoning.arXiv preprint arXiv:2505.02387, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.777215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.777215Z digest=sha256:c209f4a453e128033d6d9591c97dd39123a784a663168d75e4797f5da1b88c00

Observation 4325b17c-a6d6-41a3-89e8-0f630abe60e5 · outbound

This paper cites Christiano, Jan Leike, Tom B.

Reward Reasoning Model Christiano, Jan Leike, Tom B

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.819870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.819870Z digest=sha256:79bff308aa3d03108ed89f02f1cb5bf3f9cd1f753b5077b67233d56ce77b6e27

Observation 1c864530-7a15-455c-a7f2-e29d36148472 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reward Reasoning Model Training Verifiers to Solve Math Word Problems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.852224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.852224Z digest=sha256:81aba739f664ea67116b206de91b909fbd922dbf064c22ec0b0bdf1024e74624

Observation 99b94b53-9660-4767-8b08-eb18dde2642a · outbound

This paper cites Elo.The Rating of Chessplayers, Past and Present.

Reward Reasoning Model Elo.The Rating of Chessplayers, Past and Present

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.904944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.904944Z digest=sha256:a625f04aac027494ec09405edbbffd91ab3ae016a1baa49b388a367e2c368bc6

Observation 7e5287c2-99f5-44e1-9322-1a6eb4ea874c · outbound

This paper cites Gonzalez, and Ion Stoica.

Reward Reasoning Model Gonzalez, and Ion Stoica

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.948093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.948093Z digest=sha256:2850d089a705a1c3a04a0df49f765f4784a4622f3c9009c63dcfe30c441eb49d

Observation 5f88e40d-533c-4b66-bd39-e43d86ce9f77 · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Reward Reasoning Model On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.993826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.993826Z digest=sha256:ec5c754d34545de819b7d7f09f6ff5cf89342d76f076d1173c35523f44a16d0d

Observation 033762ca-1f77-4f47-a0ad-dbf9d7e01395 · outbound

This paper cites Scaling laws for reward model overoptimization.

Reward Reasoning Model Scaling laws for reward model overoptimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.045534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.045534Z digest=sha256:11c3dac63c90575b2713679741362a64205ca4627d2719f64c07de16bba38dbe

Observation ca1e8f79-71e1-4408-afda-e93bfc3593fb · outbound

This paper cites A Survey on LLM-as-a-Judge.

Reward Reasoning Model A Survey on LLM-as-a-Judge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.098716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.098716Z digest=sha256:c10b311b2d0e382a7dfb46af72d110c2daec614352fb20c6fcbd023f4102b1a7

Observation 9251b517-e7ab-4e26-8341-e109762f8faf · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reward Reasoning Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.142179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.142179Z digest=sha256:747571406f35b456014d1918ae2c4a4822416fbdb2d9e6816763cf7616758cdc

Observation f8aff6e2-39e4-47fe-9c9e-541080cb20fd · outbound

This paper cites Mcranker: Generating diverse criteria on-the-fly to improve pointwise llm rankers.

Reward Reasoning Model Mcranker: Generating diverse criteria on-the-fly to improve pointwise llm rankers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.178507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.178507Z digest=sha256:fbddabab126a220957cd024ebcfecee8929ed385188bbf9d54f438ce5cfba4f9

Observation 8b2953dc-7c4b-4ecd-a612-2d49d768aa75 · outbound

This paper cites Skywork open reasoner series.

Reward Reasoning Model Skywork open reasoner series

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.224822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.224822Z digest=sha256:8ef986e37077419ac275132ecd634f0df098c11193a7ce6f57bb47c758feb62f

Observation e6099a2b-3bf3-4ae4-aec6-0aec91939039 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Reward Reasoning Model Measuring mathematical problem solving with the MATH dataset

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.958482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:51.259139Z digest=sha256:2dadf8252de3d08d99d32b16900a23a9fc4a40bf08632d4b643170faf8b3bb9d

Observation f12a8b94-3cf5-4ae0-851e-f9a5556d8a5d · outbound

This paper cites GPT-4o System Card.

Reward Reasoning Model GPT-4o System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.300610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.300610Z digest=sha256:ce3f5bd9aab2ce095916ef829601cdee90720fd749953891c31c31003fc460b4

Observation 31bd0b84-c20e-4a88-b147-8b2dd8f52a19 · outbound

This paper cites OpenAI o1 System Card.

Reward Reasoning Model OpenAI o1 System Card

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.337869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.337869Z digest=sha256:1ccbcfc19a764107a47c6d39b2b3484d315ea2a3b09414325972d42c6943cdab

Observation 536d8b20-32fd-46c3-a843-3794f23bc031 · outbound

This paper cites LLM-blender: Ensembling large language models with pairwise ranking and generative fusion.

Reward Reasoning Model LLM-blender: Ensembling large language models with pairwise ranking and generative fusion

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.370094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.370094Z digest=sha256:0a51e2d9bb6803b8fa16e3eaf5e9a3403124726f0f950a13fed6a244b69c1063

Observation cd8e1793-3b22-41ac-acd8-1db43cb79403 · outbound

This paper cites SWE-bench: Can language models resolve real-world github issues? InThe Twelfth International Conference on Learning Representations, 2024.

Reward Reasoning Model SWE-bench: Can language models resolve real-world github issues? InThe Twelfth International Conference on Learning Representations, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.410454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.410454Z digest=sha256:c8e0a54565471b357d793dc4630bcde6217ba565dc3e352a0bc3e89bcfd114d7

Observation 825d3a07-6a2b-4c64-aec2-07b755b4daa0 · outbound

This paper cites Scaling Laws for Neural Language Models.

Reward Reasoning Model Scaling Laws for Neural Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.450665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.450665Z digest=sha256:696c02797611e70e2f88a65aef6739fab72f2554d7e38eb6d07c8c219eca0ccf

Observation 8afb08ed-a533-41f1-a60d-44f323f602d8 · outbound

This paper cites Maple: Multi-modal prompt learning.

Reward Reasoning Model Maple: Multi-modal prompt learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.493581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.493581Z digest=sha256:014a4e9053863ed0e0048c92f17488983a18c0b1e36c9528a016aeb5ca46e156

Observation 95386e55-d1bd-4352-9799-9053ebb635a1 · outbound

This paper cites Prometheus 2: An open source language model specialized in evaluating other language models.

Reward Reasoning Model Prometheus 2: An open source language model specialized in evaluating other language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.527785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.527785Z digest=sha256:7f900803778d1b0fadbbd6e20e48a7a165bca69f87e9df02c88d15aba74209d0

Observation ce635670-f63c-447b-afd0-6a86ceaab1fe · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Reward Reasoning Model Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.610704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.610704Z digest=sha256:e030a756cc3ea4288c5d081ab154d101a77da3f4cd26967406a65ba7966ccf7b

Observation 1ce3f9a5-3861-4ee6-b727-1c15d0e0512d · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Reward Reasoning Model RewardBench: Evaluating Reward Models for Language Modeling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.641108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.641108Z digest=sha256:a29cb3dbbfae5bfced69b3734497175d75666f6c251a39ad5e2f4a50f5bf2720

Observation d59bc306-acfb-42b5-9cef-d49b096d8c08 · outbound

This paper cites Generative judge for evaluating alignment.

Reward Reasoning Model Generative judge for evaluating alignment

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.936539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:51.675833Z digest=sha256:2996d5d7b07bd57a12fd2cd55776764374fed995111f149133a0d81ca93361db

Observation 97e93a02-a458-4cf8-bb3d-1fa1e9b98994 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Reward Reasoning Model From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.745476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.745476Z digest=sha256:51e06d4f2c01e9262d6ed2120ec1ff2d860f543bd442f29d5197b32e6d55ceb0

Observation d0dc8f84-b9f5-42c8-8c22-50f1fcd63e53 · outbound

This paper cites an unresolved cited work.

Reward Reasoning Model Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:34:55.926236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:51.708821Z digest=sha256:1be048a19ca8c5799164b0b7ff9b2cb2efa6d84b6dcf593c4f3258c542c9009f

Observation 17a41f16-cc95-4c35-847c-587a961c0c0d · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Reward Reasoning Model Rouge: A package for automatic evaluation of summaries

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.842603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.842603Z digest=sha256:ee3c8846abb5dbca121ce424f1b094256b4459bd37a699b87e273662f5c81c08

Observation 6ef38318-1a13-4659-b948-b3841db37b33 · outbound

This paper cites Let’s verify step by step.

Reward Reasoning Model Let’s verify step by step

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.791581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.791581Z digest=sha256:6f589fd3d0ee28e1db91da5d51e635855f609fa5dd6f43b4cb50d9b90c30ae51

Observation 10f0424d-d52f-4e3b-bd01-22350bb450f2 · outbound

This paper cites PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament.

Reward Reasoning Model PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.960784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.960784Z digest=sha256:8c51fd31172133be95ad9b2c9cfd3706f15a86fb918192263cdc548dd589e6b8

Observation 1944545c-20ca-4cd3-a1b0-57e105fd9972 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Reward Reasoning Model Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.910884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.910884Z digest=sha256:13ce2f156184d5dec1f3e5d532c706f0bc0562a38c342112bd87e886214ab733

Observation 0f772cdb-2858-463c-9082-4001e982e021 · outbound

This paper cites General-reasoner: Advancing llm reasoning across all domains.

Reward Reasoning Model General-reasoner: Advancing llm reasoning across all domains

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.902751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:52.052327Z digest=sha256:38be948012c734df3af7f672278cd7de7f746be4e1bd87a7e0fbf44ea65793dd

Observation fbbf5d0d-d666-4fea-beba-18ae91c32152 · outbound

This paper cites Inference-time scaling for generalist reward modeling.arXiv preprint arXiv:2504.02495, 2025.

Reward Reasoning Model Inference-time scaling for generalist reward modeling.arXiv preprint arXiv:2504.02495, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.997784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.997784Z digest=sha256:aac0040ae0f89ff322af4c04b88f5ab2a0becdc922029417cfecb4f1c8d7dbd6

Observation ca842ffc-23a6-46a5-a14a-e0cee2affeb2 · outbound

This paper cites Training language models to follow instructions with human feedback.

Reward Reasoning Model Training language models to follow instructions with human feedback

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.891726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:52.134440Z digest=sha256:65c9c4a392ef2a3d754152db27f831ffe935214444565a0ef67f1a2ff0a26766

Observation 04e333ca-9936-4e0a-b645-56707c172df2 · outbound

This paper cites Generative Reward Models.

Reward Reasoning Model Generative Reward Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.079730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.079730Z digest=sha256:4187f891829804363e0cc30ac42c382feabc4b13536ae5ba0de29f6178a3d150

Observation e156917e-cdb5-489a-9e9f-32173197839d · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Reward Reasoning Model Bleu: a method for automatic evaluation of machine translation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.199462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.199462Z digest=sha256:6fd1341c2565dafbdd6001060346e5181b7127d902662d4cd91c8267612d1a89

Observation 2e9deba9-8e08-489e-9407-3eeaba038f60 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Reward Reasoning Model Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.171617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.171617Z digest=sha256:dea39e74b9eecbac9d5d2da47c0b79c529dec92b8cff0a21e537627c683158bb

Observation d40c1b1c-1b3f-4874-8d8c-5e9cad63712b · outbound

This paper cites Qwen2.5 Technical Report.

Reward Reasoning Model Qwen2.5 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.271529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.271529Z digest=sha256:8499423af4ea428b8a1ae4b8ab5cf364f319d6490ed1e2b660513f0719f80566

Observation 7f077f48-f53e-4671-8c10-968b40c4aba1 · outbound

This paper cites OffsetBias: Leveraging debiased data for tuning evaluators.

Reward Reasoning Model OffsetBias: Leveraging debiased data for tuning evaluators

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.231093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.231093Z digest=sha256:8e7b90e8f8de5906ed8976a9220ecdfda5d8664ba25bbe83cf7df231755a6e56

Observation 6d207513-7fda-4cb8-bf09-5dedac3090c0 · outbound

This paper cites Bradley Knox, Chelsea Finn, and Scott Niekum.

Reward Reasoning Model Bradley Knox, Chelsea Finn, and Scott Niekum

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.866191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:52.342892Z digest=sha256:c4bfca5e61705333fd73c2d19427fd7bffa441e8e5e86f56ec3fee69f5c25f74

Observation 93b2115a-a6af-4994-afc3-20a88c48f4dd · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Reward Reasoning Model Direct preference optimization: Your language model is secretly a reward model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.301877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.301877Z digest=sha256:4ac6032ad829f7aa8b846fd5399676e38ddd2466eac1a03d33fa47dff40fc93e

Observation 1b28ebea-f75d-48d4-8953-78e11cbea415 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reward Reasoning Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.426178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.426178Z digest=sha256:ba348c6e9d374b0d4e458fde976cb23b25b19b70717fea4e49273aa00684fa50

Observation a5b06b04-de8e-479d-85e1-a53711d9db2d · outbound

This paper cites The limits of automatic summarisation according to rouge.

Reward Reasoning Model The limits of automatic summarisation according to rouge

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.856216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:52.383058Z digest=sha256:1cceb74d92fc50d9a6259edeeb79feee6e9d60d5afabb01a367f083d85d86a18

Observation a9d9763e-c9e5-4662-920b-503fffe91314 · outbound

This paper cites Skywork critic model se- ries.

Reward Reasoning Model Skywork critic model se- ries

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.496371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.496371Z digest=sha256:efdc5c4a80e1cdaf4a96f8768ad289df63a44007549504fb85bb670dd6171ee7

Observation c60e5510-ccfa-4c84-8520-879502e2d277 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Reward Reasoning Model HybridFlow: A Flexible and Efficient RLHF Framework

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.464502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.464502Z digest=sha256:e6068ec06b43bf88807484d784799c3b5b7ef49924c8cb8a0b0d1bece4b90b87

Observation 8540d3f5-3be1-4a25-9b67-af2314e4b64a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Reward Reasoning Model LLaMA: Open and Efficient Foundation Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.573985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.573985Z digest=sha256:b77231661399e5a7205ef16c8c6f2129ab2dff27ddbfcd33909347ee63e52c8c

Observation cfef8121-af75-407d-85ed-ec9b6453c90e · outbound

This paper cites Scaling LLM test- time compute optimally can be more effective than scaling parameters for reasoning.

Reward Reasoning Model Scaling LLM test- time compute optimally can be more effective than scaling parameters for reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.538743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.538743Z digest=sha256:b3ad56dc267003f4f9c53ab05939ddb11240ec6bb3908b013ddb624ddc0205a3

Observation a812a115-e4e8-4789-8eca-a48adb5d8948 · outbound

This paper cites Secrets of RLHF in Large Language Models Part II: Reward Modeling.

Reward Reasoning Model Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.624148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.624148Z digest=sha256:827af363072c78d53069c89f3dd38ca0918f452f534da798b4e938d7ea7bb776

Observation 9e4f17b6-25e3-441e-a397-57d7508cd05e · outbound

This paper cites Foundational autoraters: Taming large language models for better automatic evaluation.

Reward Reasoning Model Foundational autoraters: Taming large language models for better automatic evaluation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.593804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.593804Z digest=sha256:ef0b0ea6003b083a559a656b23b20f702305183985f62b9a41ca47c4d454dc93

Observation 72abce9f-4137-4ad2-b582-5787a5b1cfe9 · outbound

This paper cites Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization.

Reward Reasoning Model Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.834388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:52.701742Z digest=sha256:47e44b62006fa24178ca5bbcf2c1d54ae075dec6b9c6228a0ddf4c0162f268ac

Observation d720deb4-efae-4f39-b636-0d2064001d37 · outbound

This paper cites Self-Taught Evaluators.

Reward Reasoning Model Self-Taught Evaluators

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.667559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.667559Z digest=sha256:bc12a65a2e6788eb3996704089f182d821b7c1063f79b674663eba532d5074b9

Observation 3b22891b-1cb1-4eb9-8f68-fa2c7b066dd8 · outbound

This paper cites HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM.

Reward Reasoning Model HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.819560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.819560Z digest=sha256:89fe8b873d5ecb8e618b227818a6507983289b20f58326e7ee78b918465be7c7

Observation 341ee96e-1efe-4a1d-bfac-48453a6ff084 · outbound

This paper cites an unresolved cited work.

Reward Reasoning Model Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:34:55.823373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:52.737656Z digest=sha256:e6de2308bff675cf1d549ab9361bfd4e2c3437e98701e4fc491bcfef78ac1150

Observation 34bbcde9-445e-44e9-a73d-b4270e62c790 · outbound

This paper cites Reinforcement learning for reasoning in large language models with one training example.

Reward Reasoning Model Reinforcement learning for reasoning in large language models with one training example

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.813269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:52.773786Z digest=sha256:f06f852a7e857642f7004bbdacd61a0cc539d7e5d600714a9df9d6ffbab3a7cf

Observation ba7f0d83-0f17-420d-866d-a61a987a3d3a · outbound

This paper cites J1: Incentivizing thinking in llm-as-a-judge via reinforcement learning.arXiv preprint arXiv:2505.10320, 2025.

Reward Reasoning Model J1: Incentivizing thinking in llm-as-a-judge via reinforcement learning.arXiv preprint arXiv:2505.10320, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.940879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.940879Z digest=sha256:3da5cc7b3dd2011cfb9d2d9f7a56504f344357369955f5dc74beeeb938698b83

Observation 514a7db1-078e-4cde-b405-c3bca95521a6 · outbound

This paper cites Zhang, Makesh Narsimhan Sreedhar, and Oleksii Kuchaiev.

Reward Reasoning Model Zhang, Makesh Narsimhan Sreedhar, and Oleksii Kuchaiev

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.802137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:52.862712Z digest=sha256:37df988dc2a250a27aa44dd1f3301f96a96ccd5411d1567d3d14c217433349a4

Observation a79a40e6-edfa-40cc-b0df-2f7e4ae87bf4 · outbound

This paper cites Chi, Quoc V Le, and Denny Zhou.

Reward Reasoning Model Chi, Quoc V Le, and Denny Zhou

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.909489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.909489Z digest=sha256:58b4fc80e36cd160218181ddf808e6bfb2fdde7ab6206ac89a1b3c5f7bd0b565

Observation 6245467e-b388-4761-9d3c-06b0d0775aed · outbound

This paper cites Inference scaling laws: An empirical analysis of compute-optimal inference for LLM problem-solving.

Reward Reasoning Model Inference scaling laws: An empirical analysis of compute-optimal inference for LLM problem-solving

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.021374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.021374Z digest=sha256:5184674f7b3ae3373a81e676f3d1446a65d4635d363e6ff2bfdfd418d605a683

Observation 3914fb9b-10c8-4674-9571-b66f9dc9e801 · outbound

This paper cites Metametrics: Calibrating metrics for generation tasks using human preferences.

Reward Reasoning Model Metametrics: Calibrating metrics for generation tasks using human preferences

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.784727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:52.996374Z digest=sha256:706a925261a3d343ae9ac70cb685dc4a9ab9b0d40a6ce7e012ed3ddd6184cf60

Observation 84ed587f-2345-4098-8b28-592c52a1dbb8 · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Reward Reasoning Model Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.003907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.003907Z digest=sha256:b82955c8802567f603f6e168aa279b14fbc1f98c15f4d39369d6733f10e35c0c

Observation ed3d3685-5069-4f6d-a549-19afa6aca6af · outbound

This paper cites Learning LLM-as-a-judge for preference alignment.

Reward Reasoning Model Learning LLM-as-a-judge for preference alignment

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.759404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:53.201510Z digest=sha256:35714a3efd03ddef5b9d5f57100f70eb50e6b28a877552a235ffee8590964cd8

Observation 24c0ae7f-4b60-4b0e-9cc0-07b08648a264 · outbound

This paper cites Qwen2 Technical Report.

Reward Reasoning Model Qwen2 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.085306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.085306Z digest=sha256:1f41181bbed4170ad70a18494e8bc264ac0a0d31219223a3d56886113cb4ef4e

Observation 85e6dfe3-c4c2-442f-bd24-53cb550d790f · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Reward Reasoning Model Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.137028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.137028Z digest=sha256:c720ea98799ce0dba8af6f3c36adfa314e1e62e560c64c45b31805c8e809b163

Observation de494628-796f-48eb-8e34-1871d3509850 · outbound

This paper cites Self-rewarding language models.

Reward Reasoning Model Self-rewarding language models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.349374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.349374Z digest=sha256:126e53eeac2e45be1b5484f4c2fd582cf6661b851e1d42cd46ce67c3bb894ae5

Observation a6fdcaf4-4f08-469b-9515-792345129f24 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Reward Reasoning Model DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.259571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.259571Z digest=sha256:207714718cc601495e82fd7fdc6b352e7ffb0902050fe68ce1b10b6bbf48656e

Observation e64558c4-476e-4df2-888d-8be1b6db038b · outbound

This paper cites Self-generated critiques boost reward modeling for language models.

Reward Reasoning Model Self-generated critiques boost reward modeling for language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.740256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:53.306879Z digest=sha256:fbf1c1538a01b5d9ed7ab623c19fb4dafbdcaa5bf75248a7c59d5f9e22d031d1

Observation 3868ed14-945a-4658-936a-d99a0c28f3ae · outbound

This paper cites Gonzalez, and Ion Stoica.

Reward Reasoning Model Gonzalez, and Ion Stoica

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.509079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.509079Z digest=sha256:f83fd77c3224066bb4dcc0196d2ca10520782a25cb99764634097b15b611b3a1

Observation bf5a87e0-1d19-497d-8395-e141e26afb38 · outbound

This paper cites Generative verifiers: Reward modeling as next-token prediction.

Reward Reasoning Model Generative verifiers: Reward modeling as next-token prediction

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.714358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:53.404561Z digest=sha256:8692d5a1bbdff568775df451c10ec04645a540fd877199f3eebed9b920879bcf

Observation a5b784e3-5bd9-4f6e-949f-953ce030b548 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Reward Reasoning Model The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.466706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.466706Z digest=sha256:53584efea496ee2d4af455101fcfc83c184d99476961ea89c10bf09988c9d7a5

Observation 09dd05a2-446f-4202-8241-8fb9c98a7b3c · outbound

This paper cites JudgeLM: Fine-tuned large language mod- els are scalable judges.

Reward Reasoning Model JudgeLM: Fine-tuned large language mod- els are scalable judges

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.686012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:53.658787Z digest=sha256:5b7f00e6e0420f8ff56c3246f22d6725758119164ec7a9419c84369bf8ea765a

Observation 9df9a508-d550-4afe-bdfb-f965df1da3ef · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Reward Reasoning Model Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.545855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.545855Z digest=sha256:b7031eb303ae056890bda43b64cf6baaf02acf72cc7e7b43df63745c1eb33a20

Observation e82255d5-fafd-40d0-857b-3536af06b0d7 · outbound

This paper cites A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future.

Reward Reasoning Model A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.601769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.601769Z digest=sha256:f72bc2934fb5e4ab2990c0872ef6afd3d9542c2f828e3843a512ee687074b142

Observation f22bab45-4e75-432a-8778-e7f7e5105b1a · outbound

This paper cites Partially Adhered.

Reward Reasoning Model Partially Adhered

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.466419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:53.781321Z digest=sha256:f37100bdfb48f1f21fe4ecdf7a1719aec8a24e7d21c477944d082238bcdb0b97

Observation 5065b370-aea8-49c4-9f0b-5c6e0df169ea · outbound

This paper cites Useful but Incomplete.

Reward Reasoning Model Useful but Incomplete

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.250151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:53.828691Z digest=sha256:1a9f0b49d10b0f27e5e057affaf74d228f3c42115499421cae7cf9299fe12c81

Observation 51c7af70-e049-4117-9136-927f002999ed · outbound

This paper cites Not Detailed.

Reward Reasoning Model Not Detailed

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:54.938238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:53.874909Z digest=sha256:a23fae56856d88da50524fcf2efdbf229f23c5145ab1b462a1abc477c9673fa6

Observation 7be62865-2f07-45af-bfcd-045c54341bcd · outbound

This paper cites doi: 10.18653/v1/2024.emnlp-main.248.

Reward Reasoning Model doi: 10.18653/v1/2024.emnlp-main.248

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.573946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.573946Z digest=sha256:1fcd0c19d5175fd03a0faf2c9efda93d49a5b9bb1955ac94d9f3af4cfdcba272

Observation 6be746c7-c633-497b-844a-2731a542c577 · outbound

This paper cites \boxed{Assistant 1}.

Reward Reasoning Model \boxed{Assistant 1}

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.665928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:53.714483Z digest=sha256:a433b6696e156ed61a58c5e1deca6c8e91c330a5e743c3b532c3a8ce2bd70cd7

Pith citing papers

Observation aab485b9-07c5-43db-a0cd-505f144055de · inbound

Tournament of Prompts: Evolving LLM Instructions Through Structured Debates and Elo Ratings cites this paper.

Tournament of Prompts: Evolving LLM Instructions Through Structured Debates and Elo Ratings Reward Reasoning Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:15:37.528799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:15:37.528799Z digest=sha256:e4f439b8194c4f831aeb59ea85c60cff8c21037457421d89ebe2a0ed9a03c89f

Observation 04ec0a88-7b73-41a4-92c7-b46d6cf7ab9b · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Reward Reasoning Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.760173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:6375b9e4f084d50cf5bd3b84f9388351cc1a2d11b2e0b753512dd99c2d0d834d

Observation a2818eef-f5ce-4311-b868-782e9556d367 · inbound

GenSelect: A Generative Approach to Best-of-N cites this paper.

GenSelect: A Generative Approach to Best-of-N Reward Reasoning Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:26.527647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:26.527647Z digest=sha256:e3ef17de309021092f68101af13612562438ddad64e641d5d21406a2595d2def

Observation cf18a61e-8d83-4445-8171-7d29f8bc4780 · inbound

VRPRM: Process Reward Modeling via Visual Reasoning cites this paper.

VRPRM: Process Reward Modeling via Visual Reasoning Reward Reasoning Model

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:21:30.995972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T12:20:17.430881Z digest=sha256:27777360b6fd0fc271d5553284b35366e2ae260b7542b02ad062ba191e629062

Observation a5c71336-4ebc-41fc-83b2-a6491ddc19d0 · inbound

VRPRM: Process Reward Modeling via Visual Reasoning cites this paper.

VRPRM: Process Reward Modeling via Visual Reasoning Reward Reasoning Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T04:27:05.207869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:27:05.207869Z digest=sha256:ba539c68aa7edf2514ffb3b235d44b78b117d615fd2403535fe9f68a178ba6b3

Observation 25a4676c-98de-4856-8bc4-f4660d2fd46d · inbound

Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs cites this paper.

Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs Reward Reasoning Model

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:16:50.971419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T21:16:15.703057Z digest=sha256:c27bdbf7138d048bef1f078a0e56aad21304a1b44f48ddbc2e3307aebad07252

Observation 589bb886-b043-4291-8a39-8c5ed8c35d8c · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Reward Reasoning Model

Reference 173

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:05:31.631151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:839e9875bf179573fc806009b945432f41f6fdebfb442a3599d7de8586c0ad11

Observation c3900aa1-9feb-460d-9ceb-9b19a2dc5c8d · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reward Reasoning Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.589806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.589806Z digest=sha256:2959aa02b19c19e449e33d113fd2b2b4e0a39c73da9b46055d308ab4f181f479

Observation 5ee2d70b-a2a2-46cf-ad4e-9e71f89c5c87 · inbound

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling cites this paper.

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling Reward Reasoning Model

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:20:49.334817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T03:15:57.744706Z digest=sha256:31ad0fa82294e21c5a218093b3d3131193c322a06dcd16c58ff8208d58e0b84d

Observation c37af3d7-af75-4336-af67-24164bd603c7 · inbound

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning cites this paper.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Reward Reasoning Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:24.151892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:24.151892Z digest=sha256:fbc7c2f6839fa86893b6bc5cf5cfc15dd34448c42cf721637f00dd016f830c1e

Observation 7a625218-dee7-4b6e-b395-3d97d857f5b3 · inbound

AI Can Learn Scientific Taste cites this paper.

AI Can Learn Scientific Taste Reward Reasoning Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T18:14:50.408504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:14:50.408504Z digest=sha256:261b1e0afd2531ce68c9e7f96c6e452b0f09e97f34720c1851da12c4967426ba

Observation e63bc315-3599-4d2c-9fd4-993a8f2154b0 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Reward Reasoning Model

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.571928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:c0f3376ceed1c79162201ba02fe6ba3c9ba60505b2d1b73c550493fadc04d0b7

Observation ad125d65-8338-4ff8-b3a2-25cb045d814e · inbound

Biological Reasoning-Informed Regression for Interpretable Regulatory DNA Activity Prediction cites this paper.

Biological Reasoning-Informed Regression for Interpretable Regulatory DNA Activity Prediction Reward Reasoning Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:17:26.268285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T19:01:43.892354Z digest=sha256:ccfa4f52b1a5623d1f42fcc7d9696703077cf35f8de8601f983a66e3b4e17e38

Observation d329b7c8-0292-4432-947f-e065d7525ebf · inbound

Counsel: A Meta-Evaluation Dataset for Agentic Tasks cites this paper.

Counsel: A Meta-Evaluation Dataset for Agentic Tasks Reward Reasoning Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:49:38.204971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T14:07:59.446478Z digest=sha256:9665c441348d4b01543a16fe61874de852ef7a437f6fb820bb6bfe4eb96ad1d5

Observation 12a1c266-d55d-436a-b877-aadc29b49771 · inbound

TAPAS: Throughput-adaptive Perception for Autonomous Systems cites this paper.

TAPAS: Throughput-adaptive Perception for Autonomous Systems Reward Reasoning Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T18:24:04.760985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:24:04.760985Z digest=sha256:4c580d20db923fe3e9dec7a2d0ac6da15cda1f0e71ea4a9ebb9fbca0cca5e8c9