Pith. sign in

Paper Citation Record · LEDGER

Self-Generated Critiques Boost Reward Modeling for Language Models

As of 21 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 13 inbound Pith citation observations for arXiv:2411.16646.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16646 v3

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:58:31.026570Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T06:04:14.616880Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T12:25:43.079840Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ffe186cc-3c5b-4c67-a231-575c198eb217 · outbound

This paper cites write newline.

Self-Generated Critiques Boost Reward Modeling for Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.260257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.260257Z digest=sha256:5df8656a670851a5fcb73ba92562ed8646976b0181124e154b06c3a87051681b

Observation be74fe89-6d27-403d-90e7-8c9b03a9fafc · outbound

This paper cites GPT-4 Technical Report.

Self-Generated Critiques Boost Reward Modeling for Language Models GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.265220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.265220Z digest=sha256:b3e413ae91c101e7f937011b2f1a2f6d20701c819d173315db5d69d1cc0dca0d

Observation fdc67657-8d22-4276-bdcf-5dccef2e053b · outbound

This paper cites Nemotron-4 340B Technical Report.

Self-Generated Critiques Boost Reward Modeling for Language Models Nemotron-4 340B Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.269006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.269006Z digest=sha256:49216232f2e32bfc416c75b9143210c8f7ebb53782a243f3c08cfc95e8e5f7c0

Observation 2da54106-6976-473d-9860-52191c7b82fd · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

Self-Generated Critiques Boost Reward Modeling for Language Models On-policy distillation of language models: Learning from self-generated mistakes

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:34.017921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.274180Z digest=sha256:6ff9faac7485d4976867a5f07e7efd1ef6d788cfabacd726fe6e464ca9586df0

Observation 7692ca6f-1e13-4ddc-b1df-f6598b6d6efd · outbound

This paper cites Critique-out-Loud Reward Models.

Self-Generated Critiques Boost Reward Modeling for Language Models Critique-out-Loud Reward Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.291482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.291482Z digest=sha256:632a56d2adcb90aad708e244e7c1d7589a0dc304d1713710239757ecd023f921

Observation 5495ec06-07b3-4dc0-b3cf-22ea57457e7e · outbound

This paper cites LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks.

Self-Generated Critiques Boost Reward Modeling for Language Models LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.387291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.387291Z digest=sha256:5d1bc03711664882a56d0cdb94c7e96347036ad96d0b29a848b02ce317a0d522

Observation 10718ebf-7c38-4deb-ab5a-493340d7fa5e · outbound

This paper cites Rank analysis of incomplete block designs: I.

Self-Generated Critiques Boost Reward Modeling for Language Models Rank analysis of incomplete block designs: I

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.467193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.467193Z digest=sha256:020643b1fb25535edd039dd0053da9928d7f36e56c16f03c50c9422668e9b211

Observation d881a421-a2f7-4499-aefa-9bc0ce2211a7 · outbound

This paper cites ODIN : Disentangled reward mitigates hacking in RLHF.

Self-Generated Critiques Boost Reward Modeling for Language Models ODIN : Disentangled reward mitigates hacking in RLHF

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.981992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.471521Z digest=sha256:ce43ad33d3165db156a174ee73d1c43cc21d223a9fea9ee2c923c417ad446d10

Observation 0ed77cfe-6af7-476f-9be6-77b567048768 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Self-Generated Critiques Boost Reward Modeling for Language Models Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.475517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.475517Z digest=sha256:164d1b03765d6d573fb07a2d86510774cbde6df9899569b9f51e305aa30cf1bd

Observation f9439820-ed0a-43d8-948e-53aee259aec3 · outbound

This paper cites Reward model ensembles help mitigate overoptimization.

Self-Generated Critiques Boost Reward Modeling for Language Models Reward model ensembles help mitigate overoptimization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.962505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.479613Z digest=sha256:436ca778803c41478f16586359452679f616a233c7f7f2c999e4dd91e922cf47

Observation ebdcdafe-36ad-42ac-abc8-873f127fe448 · outbound

This paper cites ULTRAFEEDBACK : Boosting language models with scaled AI feedback.

Self-Generated Critiques Boost Reward Modeling for Language Models ULTRAFEEDBACK : Boosting language models with scaled AI feedback

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.942146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.483391Z digest=sha256:0612dab99dbe16c1764f47a49eb98b312e4c9647de60f6cb886f2e9b0ed6abf1

Observation 99a51931-811e-4536-8d67-03eec6ed9b91 · outbound

This paper cites Safe RLHF : Safe reinforcement learning from human feedback.

Self-Generated Critiques Boost Reward Modeling for Language Models Safe RLHF : Safe reinforcement learning from human feedback

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.793759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.487189Z digest=sha256:d21b0affb29d061416a1e4504a9415c11bcf43a241f39421437822179ad74234

Observation 1df86ae6-5bf9-477a-ab99-5f9324186052 · outbound

This paper cites The Llama 3 Herd of Models.

Self-Generated Critiques Boost Reward Modeling for Language Models The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.490464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.490464Z digest=sha256:accba8aea56ac257910710c559e9fc0c63d5cec21197bebf5c9bab3a1afa4c06

Observation bdd7b0af-553f-4e15-b4fe-a755fbbfd4b8 · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.

Self-Generated Critiques Boost Reward Modeling for Language Models Alpacafarm: A simulation framework for methods that learn from human feedback

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.643273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.598679Z digest=sha256:12eadfa1065b39f1d646ca0988127471a67e0000b625bc071e3507cc46b4b3cb

Observation ab384553-124d-4260-bf41-f028c093f2b9 · outbound

This paper cites Understanding dataset difficulty with V -usable information.

Self-Generated Critiques Boost Reward Modeling for Language Models Understanding dataset difficulty with V -usable information

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.621702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.608807Z digest=sha256:23f36778faf0a549663724db8cca9b619de9933e816e58a8744af0136fd2e3ba

Observation d7c37332-c054-4d80-8180-4cab9d7947a4 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Self-Generated Critiques Boost Reward Modeling for Language Models Reinforced Self-Training (ReST) for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.612122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.612122Z digest=sha256:1957fadbc4fd4c52e34d696907e162391ffa1264b1a6d73278d3745e3273eb1c

Observation b937e7ea-ae3b-41b2-839d-4f7d69ec706d · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Self-Generated Critiques Boost Reward Modeling for Language Models Measuring mathematical problem solving with the MATH dataset

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.596360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.615509Z digest=sha256:4bcb59bd34a6ae8a1e357f993733546af688bb7a5420aa0a3d8924b5653b6335

Observation 9167072b-59c6-403d-bb6a-f0beff333e73 · outbound

This paper cites Large language models are reasoning teachers.

Self-Generated Critiques Boost Reward Modeling for Language Models Large language models are reasoning teachers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.582108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.618753Z digest=sha256:7f3252418d4a170ddf8ba16cd7d2b7c346b29d544af589c4a6b7a1e1a38c71de

Observation e3d3cfd2-2f93-47f0-a5a9-4c76d2050892 · outbound

This paper cites GPT-4o System Card.

Self-Generated Critiques Boost Reward Modeling for Language Models GPT-4o System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.621988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.621988Z digest=sha256:38ecfc3f2664f61300e80af77795a5f5cbb163162f76751196bc9c1b2c3b10eb

Observation df515ede-6881-4dc0-883c-228befcbe788 · outbound

This paper cites Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback.

Self-Generated Critiques Boost Reward Modeling for Language Models Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.625399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.625399Z digest=sha256:deb1a739c73313d982875a10898c5a90256708edf72723e4e3c618453423c294

Observation ccfdbcb1-a093-42bb-882b-351bde9e5176 · outbound

This paper cites Prometheus: Inducing fine-grained evaluation capability in language models.

Self-Generated Critiques Boost Reward Modeling for Language Models Prometheus: Inducing fine-grained evaluation capability in language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.564193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.629470Z digest=sha256:365671e1a6ea24e1f5a779c7560bd6ccdb28bbf73c877332c606a804d4f25fa0

Observation 4952cc09-d83e-42c5-8f0c-60f732d50cf1 · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

Self-Generated Critiques Boost Reward Modeling for Language Models Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.634019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.634019Z digest=sha256:7d3a6940f3d888b91b2ef1be1c9b0c48ed56dcdd34ac28b58470b5a4856726c7

Observation 2b1a55d5-82df-4bee-a113-6fffd0b28b8c · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Self-Generated Critiques Boost Reward Modeling for Language Models Adam: A Method for Stochastic Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.698430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.698430Z digest=sha256:8f4379de3d9bbee17126881b9f1c103e2092769acd72a7c754102c61c105db5d

Observation 6826f4ad-82ca-4a33-befc-a9094d6f0b12 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Self-Generated Critiques Boost Reward Modeling for Language Models RewardBench: Evaluating Reward Models for Language Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.782781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.782781Z digest=sha256:2e6c46543f354d06be18dae251fb88d73c1c0fb28c679fde31571edd60cb69fe

Observation ca8cdb4b-64be-4003-824b-f6a7f005b014 · outbound

This paper cites RLAIF vs.

Self-Generated Critiques Boost Reward Modeling for Language Models RLAIF vs

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.543902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.787039Z digest=sha256:cf88c6a5a68dfaecb31281dca420d93b952f89eb6a9a25b86866c9da0958c747

Observation ead8f73b-a27e-49a0-a4c1-4f6efa6ad336 · outbound

This paper cites Generative judge for evaluating alignment.

Self-Generated Critiques Boost Reward Modeling for Language Models Generative judge for evaluating alignment

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.333152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.790638Z digest=sha256:cd3e5ac604161a84d05ab5bb3772356bc1d546d3fd2546ac56e20d9b510fe181

Observation 8b2bad33-2fdb-4f3d-a422-e48cd5c4495b · outbound

This paper cites Self-alignment with instruction backtranslation.

Self-Generated Critiques Boost Reward Modeling for Language Models Self-alignment with instruction backtranslation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.249300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.794426Z digest=sha256:91a4a3baddb18a4211793a34792690071249dc2336a3c60ff6b72ce8522a88ba

Observation 1ebf980e-f036-43f8-9ae6-cddbb62346d3 · outbound

This paper cites Alpacaeval: An automatic evaluator of instruction-following models, 2023.

Self-Generated Critiques Boost Reward Modeling for Language Models Alpacaeval: An automatic evaluator of instruction-following models, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.798012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.798012Z digest=sha256:893b7af293efd4272ca2749a66d297358694791d6794a62e4e820b251c7a7f34

Observation 0914b2ba-0d33-4926-badd-596a3b042408 · outbound

This paper cites C ritic B ench: Benchmarking LLM s for critique-correct reasoning.

Self-Generated Critiques Boost Reward Modeling for Language Models C ritic B ench: Benchmarking LLM s for critique-correct reasoning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.228047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.801204Z digest=sha256:b120b0ff18bba0e80de4b5abedb3166ac36ca646effc3d87c4156c702f6a3889

Observation 23ba04f1-cdbe-4813-9a5e-c25a3f1e13c7 · outbound

This paper cites RRM: Robust Reward Model Training Mitigates Reward Hacking.

Self-Generated Critiques Boost Reward Modeling for Language Models RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.805299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.805299Z digest=sha256:b0dd769268c4db2df738ed6d86cab58fac6546041706dd33e9fbeda1513c8e03

Observation 487d4a13-6e52-441d-a9e3-11fe3b6751c3 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Self-Generated Critiques Boost Reward Modeling for Language Models Self-refine: Iterative refinement with self-feedback

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.217052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.808678Z digest=sha256:9c7a9f13ffcb119ab5cc3c8630d0679788f1779bfa6e86b134863ea1c3709f64

Observation 9aed85a1-f604-4f0f-b5ab-0e65a84fceda · outbound

This paper cites Generative Reward Models.

Self-Generated Critiques Boost Reward Modeling for Language Models Generative Reward Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.812439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.812439Z digest=sha256:f309d1020483555c4fc86325d453fa2329c1d9c8ba556a6d32020100e8c484bf

Observation 6bba815f-51b3-421d-a58c-56362845ad24 · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

Self-Generated Critiques Boost Reward Modeling for Language Models LLM Critics Help Catch LLM Bugs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.816635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.816635Z digest=sha256:cc413dd1f5db3963dfaf88ba55cc4a52ea7b41cb8d854856bd1415f8b39ae40e

Observation 4ddc2213-a8b0-4b6e-9aea-33b298231266 · outbound

This paper cites Introducing ChatGPT , 2022.

Self-Generated Critiques Boost Reward Modeling for Language Models Introducing ChatGPT , 2022

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.206464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.824262Z digest=sha256:cdcf59700dda43ff6a419aac05803ae22c5577fd26fa010b209b689890d79141

Observation 5194ba0e-c31a-44f4-85d5-75a67c2e0cf9 · outbound

This paper cites Chatterji, Faisal Ladhak, and Tatsunori Hashimoto.

Self-Generated Critiques Boost Reward Modeling for Language Models Chatterji, Faisal Ladhak, and Tatsunori Hashimoto

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.195224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:29.871068Z digest=sha256:a804650f624de0224d527bd101e215f4ca4a9cd109fee3829a47da8c8cb788aa

Observation dff5371a-f534-47f2-8100-5053755232a0 · outbound

This paper cites Training language models to follow instructions with human feedback.

Self-Generated Critiques Boost Reward Modeling for Language Models Training language models to follow instructions with human feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.008331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.008331Z digest=sha256:93e70187cc0ff913f7ca00ae0097ea48b9a8dd93c7e261fe722c9d4057947d58

Observation f83f81e4-53fb-4b58-bc24-dc87258681b0 · outbound

This paper cites West-of-N: Synthetic Preferences for Self-Improving Reward Models.

Self-Generated Critiques Boost Reward Modeling for Language Models West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.027648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.027648Z digest=sha256:81067cef8b0f95d0a9232e354a235ddcdc5c50d4ffc948795044dc9938948d48

Observation fe18b1f5-4420-43cc-9fa1-796a0d13cab2 · outbound

This paper cites Iterative reasoning preference optimization.

Self-Generated Critiques Boost Reward Modeling for Language Models Iterative reasoning preference optimization

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.173968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:30.048582Z digest=sha256:a8b5e1efab9e231677edae91b840fe54dcc305b163f69347da47692142f3b92e

Observation 3a3ae537-c3f5-4657-af8a-3492ada82995 · outbound

This paper cites Self-Consistency Preference Optimization.

Self-Generated Critiques Boost Reward Modeling for Language Models Self-Consistency Preference Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.052464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.052464Z digest=sha256:6fb131cb9806ebfe14b2cacc02693a45e5d0d5410affbb4ba26bbe2ed46bc166

Observation 5a344517-4d59-4bf7-acb0-701819fa96b1 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Self-Generated Critiques Boost Reward Modeling for Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.056903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.056903Z digest=sha256:38836948540e252f72cb9f5a5a9bb4597400ace1eff14d6b77fbedb55c3875ec

Observation 3bae4a46-f07d-4305-adf9-30f3e4002ff6 · outbound

This paper cites Self-critiquing models for assisting human evaluators.

Self-Generated Critiques Boost Reward Modeling for Language Models Self-critiquing models for assisting human evaluators

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.061143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.061143Z digest=sha256:c0c3553e7a489fe01a188c64cef26f8867ad774a771cf4d7f965b7a664d29aca

Observation f3241364-a4d7-4f55-9ece-e712e0410b41 · outbound

This paper cites BOND: Aligning LLMs with Best-of-N Distillation.

Self-Generated Critiques Boost Reward Modeling for Language Models BOND: Aligning LLMs with Best-of-N Distillation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.107501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.107501Z digest=sha256:5bf10e605233534050d4afe9cf816dcad64952c306e11a9a8ad8c7b73e3fdd25

Observation 2073dceb-9bcd-4492-bdee-b4b33894bd74 · outbound

This paper cites Boosting Reward Model with Preference-Conditional Multi-Aspect Synthetic Data Generation.

Self-Generated Critiques Boost Reward Modeling for Language Models Boosting Reward Model with Preference-Conditional Multi-Aspect Synthetic Data Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.214477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.214477Z digest=sha256:f6108761c201f10cbeb41f4569821fb7edd9c7edd8d6e2ae11fd594f1344583e

Observation 01e1fd14-dc31-425c-8d7e-69470379eebb · outbound

This paper cites The trickle-down impact of reward inconsistency on RLHF.

Self-Generated Critiques Boost Reward Modeling for Language Models The trickle-down impact of reward inconsistency on RLHF

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.163720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:30.274814Z digest=sha256:f9cdd7155a57a8825373b6ef3188f467a916955d8775b6f848a13de34d877652

Observation 34899596-ca26-4c55-96fb-253fc2dd3352 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Self-Generated Critiques Boost Reward Modeling for Language Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.292523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.292523Z digest=sha256:1ccc40e395a737cd69e21ed79f82ccd5b0a3f5a99cecdf66f157174993449e85

Observation d9e3cab9-e24f-474d-a7aa-1fdb5f9518e3 · outbound

This paper cites an unresolved cited work.

Self-Generated Critiques Boost Reward Modeling for Language Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:58:33.151959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:30.297357Z digest=sha256:173042f793da26d62ba2dcd534b2c17bf3268d78f3a5f5c30ddb11b36ef88b3d

Observation 3646890f-9ca2-4e2c-a6b3-a73bdd7e84b4 · outbound

This paper cites Learning to summarize with human feedback.

Self-Generated Critiques Boost Reward Modeling for Language Models Learning to summarize with human feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.301418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.301418Z digest=sha256:32ff1cc0ced678a04629b2796a2c3dda978a676042d86599f114b7e38dc5d04a

Observation f3ddc1eb-74cf-4886-95f6-612dd368a724 · outbound

This paper cites Large Language Models are Inconsistent and Biased Evaluators.

Self-Generated Critiques Boost Reward Modeling for Language Models Large Language Models are Inconsistent and Biased Evaluators

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.306126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.306126Z digest=sha256:11924b2f183fdff1c166df5b5634aff3357c6a7e97825b2ca760f4f918e814fa

Observation 03829fd2-a518-4681-a276-726bdbe383ed · outbound

This paper cites SALMON : Self-alignment with instructable reward models.

Self-Generated Critiques Boost Reward Modeling for Language Models SALMON : Self-alignment with instructable reward models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:33.003980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:30.422406Z digest=sha256:ff6bb298e229869ed7a316aea0ec8a026ed1a2509822973747138721cfeb60b7

Observation 87774883-6e0d-499c-b615-1a7ab9dd2eae · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Self-Generated Critiques Boost Reward Modeling for Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.503198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.503198Z digest=sha256:cb58d7e3ffec97989d9dc6b3773450eb6c7446dfa626a823ebfa658f79a8abb4

Observation a75f3008-98fd-41c2-9d5e-f1395e529a24 · outbound

This paper cites Interpretable preferences via multi-objective reward modeling and mixture-of-experts.

Self-Generated Critiques Boost Reward Modeling for Language Models Interpretable preferences via multi-objective reward modeling and mixture-of-experts

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:32.816979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:30.520261Z digest=sha256:a539a6b3cb9a348a33c2e4b42d0a106399c0bedc0be0f45985578585fc4d2a95

Observation 68f3bf2c-00b1-44bc-a2b4-bf32d29cbcd0 · outbound

This paper cites Self-Taught Evaluators.

Self-Generated Critiques Boost Reward Modeling for Language Models Self-Taught Evaluators

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.525152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.525152Z digest=sha256:8f05258952d812a697a91574b57edf345706f3e4df50bd73c23823b4902190e2

Observation db5b1d9e-0179-4254-8595-3c1d171a551a · outbound

This paper cites Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou.

Self-Generated Critiques Boost Reward Modeling for Language Models Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:32.763179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:30.528519Z digest=sha256:012058a3d6e84e241615ea70da7c6c5ea50a172e81c7cb24ffb15e8bbdfe864d

Observation ae03238b-bd98-4a39-b8f1-4f0487bc72da · outbound

This paper cites HelpSteer2-Preference: Complementing Ratings with Preferences.

Self-Generated Critiques Boost Reward Modeling for Language Models HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.532006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.532006Z digest=sha256:e345bae4f36ef3c2826e66695119918845c4303350be7747cb26440fda6bd494

Observation 5277e0aa-39f4-4e3d-9cdb-0541f6c20deb · outbound

This paper cites HelpSteer2: Open-source dataset for training top-performing reward models.

Self-Generated Critiques Boost Reward Modeling for Language Models HelpSteer2: Open-source dataset for training top-performing reward models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.536214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.536214Z digest=sha256:dde3b820ab5a3ec0c320b9fabc883b3a44815f6f6280a8dcb0c9446b168fcab2

Observation 462bf89c-54f7-4454-8d6b-7376217d64f3 · outbound

This paper cites H elp S teer: Multi-attribute helpfulness dataset for S teer LM.

Self-Generated Critiques Boost Reward Modeling for Language Models H elp S teer: Multi-attribute helpfulness dataset for S teer LM

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:32.707233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:30.539941Z digest=sha256:1cbe34ba94e2c0ca936e20cec774cfd693a90c8d5f5bdf0e313f6c813eb55e04

Observation 3be7e336-c88e-4252-8529-764dd791fde2 · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Self-Generated Critiques Boost Reward Modeling for Language Models Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.652167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.652167Z digest=sha256:f864d5c223a0fe487793f98dd958c30e0df0822fd252c1c9447c01c2ca050e84

Observation c0a009cb-092a-4642-9c3f-d7f0457bafc0 · outbound

This paper cites Smith, Mari Ostendorf, and Hannaneh Hajishirzi.

Self-Generated Critiques Boost Reward Modeling for Language Models Smith, Mari Ostendorf, and Hannaneh Hajishirzi

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:32.482742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:30.718343Z digest=sha256:d64458cb507b09b198045f372522a553ec50022edc42a9b990a508f01191c7c3

Observation fe73cde1-93d5-4685-850c-a737b1a660a5 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

Self-Generated Critiques Boost Reward Modeling for Language Models WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.721911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.721911Z digest=sha256:7e38aa2d0bb7a3c359514a266c744cc4741fc3f48e7f6c3f2af64b2efb3f0ca5

Observation 92f1ac4a-af66-4441-a1e6-c82f7e85207d · outbound

This paper cites The Perfect Blend: Redefining RLHF with Mixture of Judges.

Self-Generated Critiques Boost Reward Modeling for Language Models The Perfect Blend: Redefining RLHF with Mixture of Judges

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.725632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.725632Z digest=sha256:390fa8036407654c5035285646bfe3944ced375ed8e9bd4c1f2ba514b7bf994a

Observation 6d441469-538a-448d-97c9-7de403a71d73 · outbound

This paper cites Predicting text preference via structured comparative reasoning.

Self-Generated Critiques Boost Reward Modeling for Language Models Predicting text preference via structured comparative reasoning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:32.471803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:30.864858Z digest=sha256:a0350a86843ccae7141321b0ec7c800a25cfafaf8a040c8cfe1932c858fdca83

Observation c085e414-9992-48f3-a9ba-74ff99df1479 · outbound

This paper cites Improving Reward Models with Synthetic Critiques.

Self-Generated Critiques Boost Reward Modeling for Language Models Improving Reward Models with Synthetic Critiques

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.870757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.870757Z digest=sha256:bbd77ac1d3720e2f70fad6238f2e5437281fdc5ac4d8200c2b33baf9f9088fa9

Observation 0bf6c7bb-daa5-409e-8f91-038ad55d8786 · outbound

This paper cites Self-rewarding language models.

Self-Generated Critiques Boost Reward Modeling for Language Models Self-rewarding language models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:32.412079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:30.876443Z digest=sha256:f9bf5b4588e0edd0ed364f6dd88a46445f874f307470afff5a1b76392d4a3559

Observation 37055a91-22f2-4359-b46f-04cdc0d4c6e6 · outbound

This paper cites ST ar: Bootstrapping reasoning with reasoning.

Self-Generated Critiques Boost Reward Modeling for Language Models ST ar: Bootstrapping reasoning with reasoning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:32.344182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:30.880625Z digest=sha256:b92397abfbcfcbead81fbdf77c8854fb62d3055cb3bdac4c24772959b1a42404

Observation bdbc6414-9384-429c-921d-2bcaafe35f5f · outbound

This paper cites Evaluating large language models at evaluating instruction following.

Self-Generated Critiques Boost Reward Modeling for Language Models Evaluating large language models at evaluating instruction following

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:32.291325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:30.943792Z digest=sha256:d0cd197a056232721979702a63e04943124db67d8b480f348296c795d6cc25f3

Observation 87930707-cc45-4c6a-b59d-14439ba24ee1 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Self-Generated Critiques Boost Reward Modeling for Language Models Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:31.011359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:31.011359Z digest=sha256:873461119d1f5cc36152d490194f715e98ad53246ec95d09d548c5b1146eb936

Observation cd504e4e-fc1e-4e42-864f-f59ca8389635 · outbound

This paper cites Gonzalez, and Ion Stoica.

Self-Generated Critiques Boost Reward Modeling for Language Models Gonzalez, and Ion Stoica

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:32.213472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:31.017046Z digest=sha256:70e73796ada1ea6344d4d205967ad9b337e5d1ba6adda456f153064535884691

Observation c5263d52-72f4-4547-965e-7db28e34b226 · outbound

This paper cites Law of the Weakest Link: Cross Capabilities of Large Language Models.

Self-Generated Critiques Boost Reward Modeling for Language Models Law of the Weakest Link: Cross Capabilities of Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:31.022253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:31.022253Z digest=sha256:29a5902fd2e9546db4ea4347b1670d15f14519063ba76a5931cedf4740a64025

Observation 9b5d11c8-4154-4cfe-b795-fb9bb11cc8c9 · outbound

This paper cites Iterative data smoothing: Mitigating reward overfitting and overoptimization in RLHF.

Self-Generated Critiques Boost Reward Modeling for Language Models Iterative data smoothing: Mitigating reward overfitting and overoptimization in RLHF

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:58:31.736219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T12:58:31.026570Z digest=sha256:941f7804f97b3fdfc2120cc2991f67bb6ee1cade9d6921bc98fc35d70bcfc6db

Pith citing papers

Observation 4d8777be-64db-4555-bd5a-c50a6fe447bf · inbound

In Context Learning and Reasoning for Symbolic Regression with Large Language Models cites this paper.

In Context Learning and Reasoning for Symbolic Regression with Large Language Models Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:53:21.144572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T18:50:40.378720Z digest=sha256:a702cf73933f0985d5826cd8fb32e878c85c808af58200ee7fe5f4692b69aba1

Observation 11b227f3-7a53-48b0-a279-df007f79eb27 · inbound

AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling cites this paper.

AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T11:43:24.850756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:43:24.850756Z digest=sha256:3e768b2f6b89a7e83d7a134f65e6d7323be20e8f42c5093901420d80dd810e14

Observation 199020c2-2a2a-4000-8eb8-0509df32500a · inbound

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs cites this paper.

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 278

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:51:29.406469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T15:51:29.022336Z digest=sha256:9c014aec0a77de2a6aa18616926e129fb55a551648f61ec5a6a9ab472d5eb6a1

Observation 860b6c67-1087-476e-b6cd-f5f4d8e47482 · inbound

CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis cites this paper.

CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T22:26:46.683078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:26:46.683078Z digest=sha256:8cb5c4ddfbea8406d7aab5cfdd5e6f951ed54b224453d9982e73e89db1b52580

Observation 2bf10488-a08d-454a-a18f-e96d366e2c8e · inbound

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information cites this paper.

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T13:25:52.136038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:25:52.136038Z digest=sha256:b9582dd36ead40ac2c61c41484b18d57ef50483f7158b64712be8875d7d6bfdd

Observation d4229abe-225d-4a77-996c-286beb3372a5 · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.481768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.481768Z digest=sha256:2b3989821f33d9f01b6cf56467d48ef1e5ee3b10dfac37e2aa4640170b23ae67

Observation e96f158f-e359-4e4f-bac7-52b63dee4c60 · inbound

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning cites this paper.

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:14.616880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:14.616880Z digest=sha256:b0312be867bf988a2dbe0dcfce5d96a2dc63ea9b39a665b27ae6772439e48384

Observation 13be2a22-8da4-46b1-9941-d6d1fb329489 · inbound

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge cites this paper.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.757228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.757228Z digest=sha256:70728366944d27d2b143462270032e3dd61942ac70156b39c0580e5f8e2efb7e

Observation faf6a9ce-0c08-4f47-8de0-a77a1dec0b40 · inbound

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models cites this paper.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.445894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.445894Z digest=sha256:49639ce3c382381e241c8f82e71c4219cad85d8b4c865435890b9a0fbaf868ed

Observation a978274c-3e04-4b41-9561-e696d770c2d0 · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.139918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.139918Z digest=sha256:afd7dac9692ddc3a6cebc99fc9f34f064b32443a5b2c04bf71fa9caca5994165

Observation 422ef1bf-506e-4adc-bdba-de513b822b03 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 282

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.294496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.294496Z digest=sha256:a930caa6b756ebac01bede2b015ab43b4a1873c61bca518af2a17116f85ff4d0

Observation 58c58590-3bd8-48e6-b25b-a0b1448d2c79 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.461763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:de061c35fec59791180f94ac9c13b6976418a5f9fb2a5a1375ff1745e1e1ef08

Observation 0dfaad6f-0025-45e8-99fb-15e8ced842ba · inbound

Test-Time Verification for Text-to-SQL via Outcome Reward Models cites this paper.

Test-Time Verification for Text-to-SQL via Outcome Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T12:25:43.081624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-01T02:10:13.533450Z digest=sha256:6f0e76ff1c595839d30c30ba9a87f170d055544f191058eed52c3eeb4dea39e7