Pith. sign in

Paper Citation Record · LEDGER

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey

As of 5 August 2026, this Paper Citation Record lists 100 of 237 outbound references and 2 inbound Pith citation observations for arXiv:2510.01925.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.01925 v3

Coverage vector

measured 100 of 237 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:50:14.455925Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:33:55.284758Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T07:34:21.193405Z

Reference resolution

100 of 237 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6a95a03b-b4f4-49b2-b6ff-8f78f760ae73 · outbound

This paper cites Sparks of artificial general intelligence: Early experiments with gpt-4,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Sparks of artificial general intelligence: Early experiments with gpt-4,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:11.446684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:11.446684Z digest=sha256:e1f2961cced579963ee82f646a92d60db7413b7ef0767897665875b828c55bdf

Observation b6e20d2d-894b-47ca-ae63-178c0d13bf54 · outbound

This paper cites A survey on medical large language models: Technology, application, trustworthiness, and future directions,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey A survey on medical large language models: Technology, application, trustworthiness, and future directions,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:11.538204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:11.538204Z digest=sha256:f2e0fdf79638608dd279a16e4162d0875cd2bb46f58ca2bc374a76e04cfd3083

Observation 65b334fc-5c3c-40d4-838f-f19870cc37c1 · outbound

This paper cites Vision-language models for vision tasks: A survey,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Vision-language models for vision tasks: A survey,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:11.629883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:11.629883Z digest=sha256:677c0e0a883e5dc212c311cb80fdfab603c8688510e9a7b7093418c9309e1cb6

Observation c370ca9e-e056-49fa-8d3e-cb3afc8a303f · outbound

This paper cites Bridging the linguistic divide: A survey on leveraging large language models for machine translation,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Bridging the linguistic divide: A survey on leveraging large language models for machine translation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:11.737873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:11.737873Z digest=sha256:0760f43b4227a09d712d016da91af23cba01e1752d363164f2de4ea6bebcaf06

Observation 1d8737de-91ae-41cd-92b0-e8f84df5295c · outbound

This paper cites A survey of large language model agents for question answering,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey A survey of large language model agents for question answering,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:11.882413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:11.882413Z digest=sha256:78ea33d2250a93249acf37d17b0cd340a627ab9bf5ba29a327cc323262a5ba80

Observation ae81b415-d7fe-4ba0-985b-9ba7ec0433a0 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Chain-of-thought prompting elicits reasoning in large language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:11.992118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:11.992118Z digest=sha256:02755f4b44176b1fe734d5e6ec69d240333c4352d8e9b285fba7616c71ebe05f

Observation 18adfcb1-2319-44cd-9f3d-1d7c1e4aa7ff · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Tree of thoughts: Deliberate problem solving with large language models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:12.143679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:12.143679Z digest=sha256:9157dee1181c0095503bffc7fb08306dfb601f7c8f5e10972c1f2be9c08fd9ab

Observation c9f8af63-0dc6-4316-b611-4a2427c02a59 · outbound

This paper cites Solving quantitative reasoning problems with language models,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Solving quantitative reasoning problems with language models,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:12.214479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:12.214479Z digest=sha256:947ed23d9de7ae84dbc20a2a0525458716cfb36471e819f1010bba2e67814b46

Observation f644d140-efad-432d-a2ca-8e7decfdc0ae · outbound

This paper cites Mint: Boosting generalization in mathematical reasoning via multi-view fine- tuning,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Mint: Boosting generalization in mathematical reasoning via multi-view fine- tuning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:12.420548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:12.420548Z digest=sha256:8fbea6fc37313a572f3f26219c57b71167ffa619178ed4a5cbdbc9d3066ead71

Observation 4c3fccd9-6d9c-47b9-813a-21767d5ae8a1 · outbound

This paper cites Robust visual question answering: Datasets, methods, and future challenges,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Robust visual question answering: Datasets, methods, and future challenges,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:12.587474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:12.587474Z digest=sha256:66a45d2a203353a6b48ceaf0b85e6ca128a460a13d54c1a35365bc955a5a5d8e

Observation f8981c89-8024-4280-9ae5-cc7fbbd06622 · outbound

This paper cites Learning from mistakes makes llm better reasoner,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Learning from mistakes makes llm better reasoner,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:12.732400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:12.732400Z digest=sha256:b6bfd7606c54f9047bbb614e66031fcb9fffc0d0702a8112b9894e37d10811d2

Observation 3a065f67-c3ad-430a-b49e-c3d5149b28ec · outbound

This paper cites Openai o1 system card,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Openai o1 system card,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:12.808295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:12.808295Z digest=sha256:517ddb8c779438944bf7e00db88f17f527f71313dcee8254198a162d1ab0689b

Observation 950a09d0-1904-48a1-9562-6f0e773b16ae · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:12.975965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:12.975965Z digest=sha256:e51a694158efb21dfb81d1171ae775f52bb0ad0b2d3ef28e129df26e4ae6453b

Observation e4325a14-3785-49b1-93f3-cf31d7f3b076 · outbound

This paper cites Tulu 3: Pushing frontiers in open language model post-training,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Tulu 3: Pushing frontiers in open language model post-training,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:13.143684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:13.143684Z digest=sha256:046d5081cdcced59709336cc5ed51f6e75358dd7a82f15bb58e8d4230ce7209c

Observation 981c5ca8-a72c-4135-ad62-cfd7f79f12f8 · outbound

This paper cites Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:13.248734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:13.248734Z digest=sha256:e636ef24039fefac1be630b83d35cc3ec0ac70367a7baad1f602166df932e6f5

Observation 4c08e556-8a2e-4a4a-940c-32915323eb0a · outbound

This paper cites Improve mathematical reasoning in language models by automated process supervision,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Improve mathematical reasoning in language models by automated process supervision,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:13.351189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:13.351189Z digest=sha256:ff4af9cebe8ce86252c25003fd00eba152c2e0ef43317c1018451501bc2a7f58

Observation 6af97ff9-afa4-4bb8-9479-d4dc555b1d68 · outbound

This paper cites Advancing process verification for large language models via tree-based preference learning,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Advancing process verification for large language models via tree-based preference learning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:13.423567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:13.423567Z digest=sha256:05f967b8dafd1df49b652ce3c1b3cd5707d78a860835ba18729596eded882ede

Observation c7e351f8-e6d2-48c5-9501-a5d1086c29ab · outbound

This paper cites Token-supervised value models for enhancing mathematical problem- solving capabilities of large language models,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Token-supervised value models for enhancing mathematical problem- solving capabilities of large language models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:13.579366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:13.579366Z digest=sha256:f66eda5edbbb985f8d7100c0027a5a33e5affc983e5d5954c3c856b52d2db780

Observation c4b5c155-a019-4f16-a769-d041a4e76f53 · outbound

This paper cites Coarse-to-fine process reward modeling for mathematical reasoning,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Coarse-to-fine process reward modeling for mathematical reasoning,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:13.720766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:13.720766Z digest=sha256:cee37971edc211da776535d116ebfe10333e547d2bf2155018641e56bbc4051b

Observation eca445c4-b06e-4aa0-b5bc-31c18bc3b413 · outbound

This paper cites Visualprm: An effective process reward model for multimodal reasoning,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Visualprm: An effective process reward model for multimodal reasoning,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:13.786433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:13.786433Z digest=sha256:08053dee31de1e86fdd1910da3838c12f68c1ffd2c6d5d42958fc6bf6520ed24

Observation 5443f8b8-3ebd-443c-9a89-70cf63c28350 · outbound

This paper cites Towards hierarchical multi-step reward models for enhanced reasoning in large language models,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Towards hierarchical multi-step reward models for enhanced reasoning in large language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:13.846001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:13.846001Z digest=sha256:ccc73caa09dc334d21dd7bcb7d1d1d7011b02426e6820e0729cb033571629cc2

Observation d2640117-11e2-4710-adc8-743fb1fe0347 · outbound

This paper cites Adaptivestep: Automatically dividing reasoning step through model confidence,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Adaptivestep: Automatically dividing reasoning step through model confidence,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:13.909596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:13.909596Z digest=sha256:339bf18f0f65dfa5982da22b36c24b84220f4ea9ca30a995476c9bd7a09f74d6

Observation 28d0a2a0-ff23-4c26-89cd-1a4483651a9a · outbound

This paper cites Vilbench: A suite for vision-language process reward modeling,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Vilbench: A suite for vision-language process reward modeling,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:13.958049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:13.958049Z digest=sha256:3edfb08107c96a9bc82ca2cdbf46bd762e99d2bd50d3bda2cb7952fc713c47b3

Observation ba058eca-24e4-461c-b5fb-cf29e90f5df4 · outbound

This paper cites Retrieval-augmented process reward model for generalizable mathematical reasoning,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Retrieval-augmented process reward model for generalizable mathematical reasoning,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.001961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.001961Z digest=sha256:41076dc94d3cceee8e1386cbb4217cee3453e7cca345f71614701dc3ba60a75e

Observation 8ccec1ab-d979-453c-a289-11b26413f3ca · outbound

This paper cites Making large language models better reasoners with step-aware verifier,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Making large language models better reasoners with step-aware verifier,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.007078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.007078Z digest=sha256:469de94d1d455c434874c298f0aa3041da9d024db10b92598387e60f72db90c5

Observation 6a8f422e-487f-46a6-9b20-be7b6bc2fd49 · outbound

This paper cites OVM, outcome-supervised value models for planning in mathematical reasoning,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey OVM, outcome-supervised value models for planning in mathematical reasoning,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.011843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.011843Z digest=sha256:9aac69c48eaaf43af58c096afc4b038eb81478edee0cf035c164a2abb39284b7

Observation 7b64f80e-d236-46b9-b8cc-72d92c42beea · outbound

This paper cites Let’s verify step by step,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Let’s verify step by step,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.016975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.016975Z digest=sha256:1901868bfcf579ad64ee28b548387aca4ddea27eb635117bd69a87980a689b92

Observation 433a7e1e-6775-4291-9644-ea8689f8ee86 · outbound

This paper cites Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.022888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.022888Z digest=sha256:9677e2c33a661115f08111c649a1de101a16bd54e254594bc2ab7f8080c8dcc7

Observation 7a3c37f8-591c-4326-a71f-cb7f409997ba · outbound

This paper cites Multi-step problem solving through a verifier: An empirical analysis on model- induced process supervision,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Multi-step problem solving through a verifier: An empirical analysis on model- induced process supervision,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.027754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.027754Z digest=sha256:68b5fbb50ee07f56298ea0618f0a1a3c7f1cd3b5afb7b73beb6b93d9f6849dd7

Observation 2f72ca02-c683-4cf8-98bd-f5772585239a · outbound

This paper cites Glore: When, where, and how to improve llm reasoning via global and local refinements,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Glore: When, where, and how to improve llm reasoning via global and local refinements,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.033978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.033978Z digest=sha256:b160b9829e75a045b83e8fd029f54ba7bfbbf4ca7580a3d3d0b19c899469027d

Observation 88c18d37-aa41-4b9e-bb90-69f189371aeb · outbound

This paper cites Autopsv: Automated process-supervised verifier,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Autopsv: Automated process-supervised verifier,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.040227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.040227Z digest=sha256:77412eb51c88a861284d0048af98c5d5c3d2f1ba944defd54d9bade8e8c930a7

Observation 8968b538-1925-4c66-aa34-f30e9132d774 · outbound

This paper cites Rewarding progress: Scaling automated process verifiers for llm reasoning,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Rewarding progress: Scaling automated process verifiers for llm reasoning,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.045182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.045182Z digest=sha256:98c532696fdccf9cd37031bd0d40f66ec8254ea00d8079bb8d77540bf73ff76c

Observation b7ab4a52-cfa4-4aa9-a172-a1ab6f8fd92a · outbound

This paper cites Entropy-regularized process reward model,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Entropy-regularized process reward model,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.049808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.049808Z digest=sha256:7bd9d5cb9cbd6ff31307efe5a21a5357659011100e30fbce2f2e5b16e3eda32a

Observation c77ee0f3-2e86-47b5-8f74-a983dbd36d53 · outbound

This paper cites The lessons of developing process reward models in mathematical reasoning,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey The lessons of developing process reward models in mathematical reasoning,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.055627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.055627Z digest=sha256:f022a60a99f8dff4c51a308163bba996ce0051008f509c986aa28bf19b68cd40

Observation ab145006-c49c-4cf3-b291-edabb76153e2 · outbound

This paper cites Athena: Enhancing multimodal reasoning with data-efficient process reward models,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Athena: Enhancing multimodal reasoning with data-efficient process reward models,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.060728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.060728Z digest=sha256:2905c992bd2f7471065965c38fa0f22578482c6b3091fe08d52fec5015b4002f

Observation befc1ed4-b622-4123-bab2-c94636ece7cb · outbound

This paper cites Reasonflux- prm: Trajectory-aware prms for long chain-of-thought reasoning in llms,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Reasonflux- prm: Trajectory-aware prms for long chain-of-thought reasoning in llms,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.065811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.065811Z digest=sha256:ddc6a9e49f4c8a25f0da105dcfa9637d5c9df6d6ebbdb29299918d90b539f796

Observation 4408501a-b69e-4981-9330-cf4d776901e0 · outbound

This paper cites Better process supervision with bi-directional rewarding signals,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Better process supervision with bi-directional rewarding signals,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.070861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.070861Z digest=sha256:1781d84a10abb87eab59e11bcfb35195d0090fdf4442cf4a5671d1ee266860bb

Observation 8dddc48c-0f60-4287-9611-42325f6876af · outbound

This paper cites Duashepherd: Integrating stepwise correctness and potential rewards for mathematical reasoning,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Duashepherd: Integrating stepwise correctness and potential rewards for mathematical reasoning,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.076069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.076069Z digest=sha256:222ac010979077d9d06c10415e1dfb38967c7fe1f85b5d8127eaa3768ecc9850

Observation f3afe761-2cb0-4162-91ac-f4de9e071fab · outbound

This paper cites Llm critics help catch bugs in mathematics: Towards a better mathematical verifier with natural language feedback,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Llm critics help catch bugs in mathematics: Towards a better mathematical verifier with natural language feedback,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.081769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.081769Z digest=sha256:e25b02c25ecab6b66e8c6e17757e6af9b4639a106c7517398af4b96a27675768

Observation ccae60fd-d920-45d0-9365-716493bd0814 · outbound

This paper cites Verifierq: Enhancing llm test time compute with q-learning-based verifiers,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Verifierq: Enhancing llm test time compute with q-learning-based verifiers,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.087042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.087042Z digest=sha256:c6cb5229123d7ee59a93df9c1c26dc114a1ee74713e7791efccf3d6fb17c198b

Observation 27ef61a6-9cc6-40ba-b6a0-f5ccb540c2ca · outbound

This paper cites Process reward model with q-value rankings,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Process reward model with q-value rankings,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.091680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.091680Z digest=sha256:91ccd84e324b0ccdbd810b19326cd755a2c4d642041eab222a3607673c6517e3

Observation f7882463-865e-40ff-ac78-288d7aadc790 · outbound

This paper cites Free process rewards without process labels,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Free process rewards without process labels,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.097267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.097267Z digest=sha256:cc5e1d73f24d1d8a5f3c15b4a6a795e2a8ed71e61ba9149c8aa87a3f0597dd15

Observation 01adce31-8d2d-4fa5-884e-7fe0bc7392d0 · outbound

This paper cites Tdrm: Smooth reward models with temporal difference for llm rl and inference,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Tdrm: Smooth reward models with temporal difference for llm rl and inference,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.102294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.102294Z digest=sha256:887badca5477206e2c23b648fef0fa03651779a4d256c9140fe45403435d2115

Observation a200b818-8382-436a-a452-8143c0dff7a4 · outbound

This paper cites Cold: Counterfactually-guided length debiasing for process reward models,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Cold: Counterfactually-guided length debiasing for process reward models,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.107620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.107620Z digest=sha256:f9d4dc8a8afa109ef421df0a32c0234e5e76c04c94a3f38fb96fda723201a284

Observation e89cb355-f8ea-4ae0-b6a0-ef27e848e98c · outbound

This paper cites Judging LLM-as-a-judge with MT-bench and chatbot arena,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Judging LLM-as-a-judge with MT-bench and chatbot arena,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.112528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.112528Z digest=sha256:87ab0e12b8a2c8d286938110564583dc13b269fba194b0e17f47ffc5e4ef0508

Observation 9d2ee219-2bc4-4091-829f-2b661ddf8888 · outbound

This paper cites R-prm: Reasoning-driven process reward modeling,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey R-prm: Reasoning-driven process reward modeling,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.117447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.117447Z digest=sha256:52a713bb6ae981ccaddbc3cdf7bd5eab555a2fde13979543e976f83b17710f81

Observation be1a5fa3-5341-4c3e-9234-ef468524ca33 · outbound

This paper cites Genprm: Scaling test-time compute of process reward models via generative reasoning,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Genprm: Scaling test-time compute of process reward models via generative reasoning,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.122164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.122164Z digest=sha256:ab7a16426a420ce4dd0e297bbd355fba300f179e993803e478a38765108d8b95

Observation 4d51a8c8-48d9-4934-809c-7119f83912af · outbound

This paper cites Scaling evaluation-time compute with reasoning models as process evaluators,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Scaling evaluation-time compute with reasoning models as process evaluators,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.126809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.126809Z digest=sha256:b7799e501c31c00ef6266c96be6b0f879fe42718c0afdd9acdf93de3531df650

Observation 76d66861-1957-4ba2-b2fe-dedda3c8cbd9 · outbound

This paper cites Spc: Evolving self-play critic via adversarial games for llm reasoning,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Spc: Evolving self-play critic via adversarial games for llm reasoning,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.131449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.131449Z digest=sha256:4f76c9bf0891772718c7278c3c260fec2aefadd35361b695b1353234829d8294

Observation bbe06494-acc8-4c38-a858-43c3b4aad374 · outbound

This paper cites Process reward models that think,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Process reward models that think,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.136042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.136042Z digest=sha256:c8eae851281212b301dffbecc2f50db4a3f4a7808f963c086a5a79786ac9cd17

Observation 61214e6f-b510-4091-b14d-ee68f711d656 · outbound

This paper cites Stepwiser: Stepwise generative judges for wiser reasoning,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Stepwiser: Stepwise generative judges for wiser reasoning,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.140823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.140823Z digest=sha256:9ffadfddc3fd938b95d528562a7c6b0bab6b6394f73bf169aa6eefdef59a2ffc

Observation b06b301d-a20d-49f7-9737-a6f4102106fc · outbound

This paper cites Solving math word problems with process- and outcome-based feedback,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Solving math word problems with process- and outcome-based feedback,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.145897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.145897Z digest=sha256:99ae66da385547daf041f454c93712e24de18bada470a9785e0080e914e292c0

Observation 1cbbbe4a-bad6-4e2f-8364-245db7f3694c · outbound

This paper cites Training verifiers to solve math word problems,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Training verifiers to solve math word problems,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.151147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.151147Z digest=sha256:c018879c272323589d9b43b95eea0a895ffc59cc597c59685a9d3662c3051211

Observation 76ad99b6-4c7c-45fa-8c2e-99022405ebaf · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Direct preference optimization: Your language model is secretly a reward model,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.156123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.156123Z digest=sha256:8aaa43bfc4f2e971d12848e48e8198138b6876ba1236232aab8c4d153b43d4d2

Observation 24752388-79e7-4524-a446-463b178d4f31 · outbound

This paper cites Inference-time scaling for generalist reward modeling,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Inference-time scaling for generalist reward modeling,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.161725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.161725Z digest=sha256:1e56d0cfea23fd2adef427349b946ee32777a3694d7823bd23eec9a75f1f5dc8

Observation 0c6389cd-8e02-4b3d-9cbf-f15607c33540 · outbound

This paper cites Rm-r1: Reward modeling as reasoning,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Rm-r1: Reward modeling as reasoning,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.166927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.166927Z digest=sha256:d1f905317b6170173f232aa5ad14487707badf54aa3480b0202655af31e691bc

Observation cdee736f-d475-4b72-899c-0a4c02b6b925 · outbound

This paper cites Ticking all the boxes: Generated checklists improve llm evaluation and generation,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Ticking all the boxes: Generated checklists improve llm evaluation and generation,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.172485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.172485Z digest=sha256:c4b9e89d87a01ce202746678c77cb781afe778c0f71fc67e5c4733b032f5ce1c

Observation e4998b6c-4cd6-40d7-a80c-5d98d63c8ce0 · outbound

This paper cites Generative verifiers: Reward modeling as next-token prediction,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Generative verifiers: Reward modeling as next-token prediction,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.179185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.179185Z digest=sha256:8ffb4030db1bd57f46b20037a790d318f3267e32da18235be35d259cf3a8c568

Observation 2ebd94fd-e49a-4f0e-95f8-147d6e6895dc · outbound

This paper cites Critique- out-loud reward models,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Critique- out-loud reward models,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.191895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.191895Z digest=sha256:e5602765c582d2a3c1dcd3a03db6cb2579c32bd14fbb15f8cf84ce4a1fa591e6

Observation 20c3961f-533c-47ad-86e2-49c113d7c270 · outbound

This paper cites Learning to reason for factuality,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Learning to reason for factuality,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.198834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.198834Z digest=sha256:d8e592d0f74ca430dff95190083cb5be24a6983ee52b7d06b964612cd864d873

Observation fd9daef2-a439-4c94-b2ae-f505cb3dd325 · outbound

This paper cites Internlm2 technical report,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Internlm2 technical report,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.204462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.204462Z digest=sha256:833c8626bc8d45d8cd691e9be38afaf4d0f6f439ccbfb468f9c20a9f9a6cf823

Observation 61b89141-ee5b-4567-99f3-70931fcdeeca · outbound

This paper cites Advancing llm reasoning generalists with preference trees,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Advancing llm reasoning generalists with preference trees,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.210688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.210688Z digest=sha256:967c0cd482a54e72ecc72d72031f6d6367b8c00af2c6b5b3c51f2c20c5bc34aa

Observation d1c2da0e-e220-400c-be60-df02b0bac22f · outbound

This paper cites Interpretable preferences via multi-objective reward modeling and mixture-of-experts,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Interpretable preferences via multi-objective reward modeling and mixture-of-experts,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.217014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.217014Z digest=sha256:1a237d52630a61b11702ea0e66690a5e14a4303cd69bf1a5d9002733b092494f

Observation 5bb0ab0c-1d27-43c0-8f6b-8918b146373f · outbound

This paper cites Llm-blender: Ensembling large language models with pairwise ranking and generative fusion,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Llm-blender: Ensembling large language models with pairwise ranking and generative fusion,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.222461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.222461Z digest=sha256:26782e1f40ae3da84c6be31befea6d38b6f3470eea92f7427cbecced8c336283

Observation 56d9afaa-df45-409e-a399-24120d42a894 · outbound

This paper cites Helpsteer2-preference: Complementing ratings with preferences,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Helpsteer2-preference: Complementing ratings with preferences,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.227834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.227834Z digest=sha256:dd47eebf7d4f56d60846fec2d93789325b1c4623acf011551215f636827b2d09

Observation 329c4246-a507-4fd1-bd8b-e85ec45be23b · outbound

This paper cites Kto: Model alignment as prospect theoretic optimization,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Kto: Model alignment as prospect theoretic optimization,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.234159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.234159Z digest=sha256:58324ce455eb1878f3cd225612a0da5028192a0dfe0be4d5593f9848eff6890b

Observation 5ef321d4-55b3-4ec8-b7ad-18158bb74f12 · outbound

This paper cites Bootstrapping language models with dpo implicit rewards,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Bootstrapping language models with dpo implicit rewards,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.240277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.240277Z digest=sha256:80e47d797237048d31ffa7d8fc25cf06fa7f4c2288d3550c1298c5c96bf99dc5

Observation 18a0ae5c-461a-4b89-83e2-6d6936083b15 · outbound

This paper cites Generative judge for evaluating alignment,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Generative judge for evaluating alignment,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.248424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.248424Z digest=sha256:aa1426e28e856750a87d42dc16b83cadec713f6451d937dd58ae7152e801611f

Observation f2687b2a-a7a6-4c83-bd06-33b5a89ad6a0 · outbound

This paper cites Prometheus 2: An open source language model specialized in evaluating other language models,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Prometheus 2: An open source language model specialized in evaluating other language models,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.253828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.253828Z digest=sha256:937b731358047e85a277dec404cced0747f72256871c012ff5a12a1d3bbdd465

Observation 306e1539-27cc-4885-90f3-3278a3c43d64 · outbound

This paper cites Foundational autoraters: Taming large language models for better automatic evaluation,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Foundational autoraters: Taming large language models for better automatic evaluation,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.261254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.261254Z digest=sha256:747663174cc05935b81397521a3c14381338eabb73c4355ae40e48d33079f368

Observation 20b41b39-442a-4110-99ec-3c25ceeed015 · outbound

This paper cites Compassjudger-1: All-in-one judge model helps model evaluation and evolution,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Compassjudger-1: All-in-one judge model helps model evaluation and evolution,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.267067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.267067Z digest=sha256:cf7d5faa442e2dd522886c0c31a66fdb5e13c262f2d39ecc6d96fd72fb9b21c9

Observation 8868fb80-26d3-4956-bc21-50e38fb58ec1 · outbound

This paper cites Learning LLM-as-a-judge for preference alignment,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Learning LLM-as-a-judge for preference alignment,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.271952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.271952Z digest=sha256:a56c94f278540563420e35c68e0224aa347908182be208e58f205b9d78ac7237

Observation 1a72b96d-905d-40dd-a5be-ff22e6f7351f · outbound

This paper cites Atla selene mini: A general purpose evaluation model,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Atla selene mini: A general purpose evaluation model,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.277964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.277964Z digest=sha256:e08b0f0e2edcf5fff66a5d845ca75c38cb06f21d00b67de5d010c94183715454

Observation 7773e415-d4c1-4ebd-a5ac-f9ea264c205b · outbound

This paper cites One token to fool llm-as-a-judge,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey One token to fool llm-as-a-judge,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.284903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.284903Z digest=sha256:27e641416474464b72dcbf9b5c95f79a8d6c46a01ef5c084e6b1c525c204689b

Observation 55d12648-5a30-45a0-84dd-349b799af72f · outbound

This paper cites Judgelrm: Large reasoning models as a judge,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Judgelrm: Large reasoning models as a judge,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.290921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.290921Z digest=sha256:6652a93ae6fe5bf3f8a692053490078392473382bc083d8fba835ced73ab5471

Observation 0db24c57-7233-4313-a0a8-b13c38ba42b9 · outbound

This paper cites Unified multimodal chain-of-thought reward model through reinforcement fine- tuning,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Unified multimodal chain-of-thought reward model through reinforcement fine- tuning,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.296450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.296450Z digest=sha256:de37f35a2f7c1855d1aaacde8a6ccb8fc63bd403bead8a247d6ca05cd104e093

Observation 133d08d4-53ce-48e3-950a-6e7d79f1f8bd · outbound

This paper cites Pairjudge rm: Perform best-of-n sampling with knockout tournament,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Pairjudge rm: Perform best-of-n sampling with knockout tournament,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.309143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.309143Z digest=sha256:7f266e6d62cb8414190a59d2e9c74613a815afd8467bda552bfd473a9e46a4ee

Observation 1e452d3f-45ed-45b5-8aa4-ac182bc87537 · outbound

This paper cites Rewardbench: Evaluating reward models for language modeling,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Rewardbench: Evaluating reward models for language modeling,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.317780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.317780Z digest=sha256:7eee6221e69fa5fa88167d8c4bee61c8ce7e909d2f75348538011d3b943065ea

Observation 43b14b7c-9b6d-4ef6-a48c-f750987d67b9 · outbound

This paper cites Rm-bench: Benchmarking reward models of language models with subtlety and style,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Rm-bench: Benchmarking reward models of language models with subtlety and style,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.324601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.324601Z digest=sha256:0ff0124359894cd776639a085339bea02534aa28bbffc92677663a2e82a2a26d

Observation 69e3e842-0aee-4b68-84a8-fc92d0a0facb · outbound

This paper cites Rmb: Comprehensively benchmarking reward models in llm alignment,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Rmb: Comprehensively benchmarking reward models in llm alignment,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.331146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.331146Z digest=sha256:6f5d87118ae2874b1dbdb34abb5f5e01badfca6a4b67b353e002b65ac270c69b

Observation 49e62bec-280d-41f1-98dd-239663fcc6e5 · outbound

This paper cites How to evaluate reward models for rlhf,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey How to evaluate reward models for rlhf,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.343788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.343788Z digest=sha256:20b78d0927aef0f594320180f51a9d875f8ac1f8111590a904c50a1cd3c814ed

Observation f519490e-5159-4253-a35f-fc5254805e0c · outbound

This paper cites Rag- rewardbench: Benchmarking reward models in retrieval augmented generation for preference alignment,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Rag- rewardbench: Benchmarking reward models in retrieval augmented generation for preference alignment,

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.349540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.349540Z digest=sha256:7e2ea40023c198d425a9b7ddcde23857a442ed3daf1524b42566c0b7b71bcb83

Observation 947b5bbd-023e-4e6a-bb0f-571c89f963fe · outbound

This paper cites Acemath: Advancing frontier math reasoning with post-training and reward modeling,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Acemath: Advancing frontier math reasoning with post-training and reward modeling,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.355432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.355432Z digest=sha256:5745bf8108731b928a4ea7ff48782b441ad6fa850793f3ccc3fc1547c63da94e

Observation ce2ecd00-6516-4a46-83ef-2014b44a5746 · outbound

This paper cites M- rewardbench: Evaluating reward models in multilingual settings,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey M- rewardbench: Evaluating reward models in multilingual settings,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.361135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.361135Z digest=sha256:7eb9825f2715959785b0b73dc2f009d783edea0b849d84e91b3b28fe9c9eef30

Observation 59035598-8f65-40ed-8ee0-8dd85b69ceaf · outbound

This paper cites Rewardbench 2: Advancing reward model evaluation,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Rewardbench 2: Advancing reward model evaluation,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.366222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.366222Z digest=sha256:39f24b86e1cf8f787943d898c57c20226ece6794a78b7f27bc0da4661504d857

Observation d86c766f-4462-41ee-88f6-e5818aa20c5b · outbound

This paper cites Rewardanything: Generalizable principle- following reward models,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Rewardanything: Generalizable principle- following reward models,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.371160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.371160Z digest=sha256:09975794170c069b9b4d30becb8086640ed8580001d133337535cfce08dd89e4

Observation f59b2732-3139-456d-be03-f4a8ca5154e8 · outbound

This paper cites Posterior-grpo: Rewarding reasoning processes in code generation,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Posterior-grpo: Rewarding reasoning processes in code generation,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.375511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.375511Z digest=sha256:7d3b34a6076cf65bf40a35e1679c96a206afcbf279491280714e97db848950a1

Observation 98ac61df-440a-452c-8999-ddd09b4dab47 · outbound

This paper cites Processbench: Identifying process errors in mathematical reasoning,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Processbench: Identifying process errors in mathematical reasoning,

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.380163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.380163Z digest=sha256:8beb7771f7dd19823b4d3ad4b4f4d99bb456335f5ad45e1fad70998e4cbcef3a

Observation 8eb11bf9-b177-4287-a981-1183e182a6c9 · outbound

This paper cites Aurora:automated training framework of universal process reward models via ensemble prompting and reverse verification,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Aurora:automated training framework of universal process reward models via ensemble prompting and reverse verification,

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.385123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.385123Z digest=sha256:d01196e7fc43ceedf184248d56d9517274234765800297a41c739f3a3409f8ec

Observation 6c6025d9-fb2c-45f5-99ab-8bc11df2feb9 · outbound

This paper cites Prmbench: A fine- grained and challenging benchmark for process-level reward models,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Prmbench: A fine- grained and challenging benchmark for process-level reward models,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.390411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.390411Z digest=sha256:aad894516fbaa252d5cc73bb9784362360f445b41ea07a899be2f804f998487b

Observation 9ce7865c-1af6-446e-9b99-b647bb8ddda7 · outbound

This paper cites Evaluating judges as evaluators: The jetts benchmark of llm-as-judges as test-time scaling evaluators,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Evaluating judges as evaluators: The jetts benchmark of llm-as-judges as test-time scaling evaluators,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.395864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.395864Z digest=sha256:ff348a187b79f531f2224777f2502e8c4dd051263632cee993e1c91c5d6507cb

Observation 139f522b-18ca-479b-bc79-008624b82976 · outbound

This paper cites Mr-gsm8k: A meta- reasoning benchmark for large language model evaluation,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Mr-gsm8k: A meta- reasoning benchmark for large language model evaluation,

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.401123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.401123Z digest=sha256:8792ad42e8ff1c445f0e0052b337765a6994115cae2059b358e8027a5f319d04

Observation 57c68f92-6f8d-45c0-a2fa-1c2e9f7f9151 · outbound

This paper cites Mr-ben: A meta-reasoning benchmark for evaluating system-2 thinking in llms,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Mr-ben: A meta-reasoning benchmark for evaluating system-2 thinking in llms,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.406644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.406644Z digest=sha256:69a07f0b52a68c5deeaf4822a2335a90ecf40964734ba4532ef6bcd313bf2da8

Observation afe1e143-9782-46d8-aded-3038aa86de8a · outbound

This paper cites Vlrewardbench: A challenging benchmark for vision-language generative reward models,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Vlrewardbench: A challenging benchmark for vision-language generative reward models,

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.412524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.412524Z digest=sha256:371357db37b3072fe7b78a281c4dffffad40f2c31e58777bbc96a7cec0d45072

Observation c61418bc-de68-446b-ac3b-90b3fccaddfe · outbound

This paper cites Mj-bench: Is your multimodal reward model really a good judge for text-to-image generation?.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Mj-bench: Is your multimodal reward model really a good judge for text-to-image generation?

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.418132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.418132Z digest=sha256:1421f690ab7e9fee2ac5b53b72060173007f9cc4208553fdac9acbfe4d278e32

Observation 12422c0f-c1ef-42dd-80c1-1f91e349b1f4 · outbound

This paper cites Multimodal rewardbench: Holistic evaluation of reward models for vision language models,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Multimodal rewardbench: Holistic evaluation of reward models for vision language models,

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.427592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.427592Z digest=sha256:555a30ebcefa8a357a4d7c1e42fdd4f83921b767f0918f6cb68ef0a21bae54f9

Observation 86227700-cb26-42e2-b957-94a3b60b3d8d · outbound

This paper cites Vlrmbench: A comprehensive and challenging benchmark for vision-language reward models,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Vlrmbench: A comprehensive and challenging benchmark for vision-language reward models,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.434641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.434641Z digest=sha256:2bf5611b63d5c58d64c0756b9fc5fc83925df0f322a9ccc758f39dec6c026845

Observation 278fd8a1-8180-483e-b0cb-8e983c23cd0c · outbound

This paper cites Large language monkeys: Scaling inference compute with repeated sampling,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Large language monkeys: Scaling inference compute with repeated sampling,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.441167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.441167Z digest=sha256:c092bad96a3d93e53818b889996f802f65bbb5a08e70ba47e14a78660d046535

Observation 6b17e9bf-f2fa-452b-969d-454cf3da10cd · outbound

This paper cites Scaling llm test-time compute optimally can be more effective than scaling model parameters,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Scaling llm test-time compute optimally can be more effective than scaling model parameters,

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.448648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.448648Z digest=sha256:f1ce14c43c6ec984a877cf568df3c6e44d210ac8a9773d8947087fba983f6c8e

Observation 7529c14e-8412-4d55-b2c0-4501c2f7866d · outbound

This paper cites Sample, don’t search: Rethinking test-time alignment for language models,.

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey Sample, don’t search: Rethinking test-time alignment for language models,

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:14.455925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:14.455925Z digest=sha256:d005ffe5bec22b0cf5562f7c5540ab93607ef5a7929af7e9243682dd341ec876

Pith citing papers

Observation f7a80c14-a5b8-4864-83fe-7e0763e01262 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-08-04T02:39:30.770950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:aacae6a783cd8778638ef802b05906edae35c6938070e09536537fd691011425

Observation 164d8e26-45ef-4cf3-8df2-4949f61a554c · inbound

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners cites this paper.

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-08-04T02:39:30.770950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T07:33:55.284758Z digest=sha256:2f5ceb68e914054af96d83d8e3c1e0f53fa3ca72d8500d2f54809ef776068572