Pith. sign in

Paper Citation Record · LEDGER

Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2312.09244.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.09244 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:53.635786Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f9169d0a-e4b5-4e96-b42f-62d01dd8571c · inbound

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models cites this paper.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.635786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.635786Z digest=sha256:887fc9948ab1a87b9e84341d730c806d031c0b57c789c082e08d3561cf8e8a9c

Observation fbe83929-5b8f-41df-8304-74f43f8010da · inbound

T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation cites this paper.

T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:26.655153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:26.655153Z digest=sha256:2566f17daa1aff1c328f5803a0a580c3f2adbd915adc81b9de1bbbb24218947f

Observation 39000e2c-02cc-43e3-bb45-c39de069dde3 · inbound

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models cites this paper.

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:30.866015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:30.866015Z digest=sha256:cafcb2bc60ae8cc8b368cf986ca053f8a20727c7ef062ae15e82ddddc831be93

Observation 85ccfc5b-4c63-4600-9e91-6f20604f8e6a · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.764200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.764200Z digest=sha256:ae8772aaf6958a69a6cb0329127b94bb1e070717d6e467a591dc6615a71a2a05

Observation 3caeb390-76d5-4292-abbc-7d1102c841bb · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.833906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.833906Z digest=sha256:9e50b0d74f58679238e6387d2a0d7b69392890a53ec372e0da1af6d6f799f301

Observation 0b8325f5-f3fe-49db-925f-63231a3714a4 · inbound

Activation Reward Models for Few-Shot Model Alignment cites this paper.

Activation Reward Models for Few-Shot Model Alignment Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.365297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.365297Z digest=sha256:b66b2581330b8ddadbbd859dbf96bd1ca998c6724eb0424283249abe316c9262

Observation 3366544c-778c-497b-a6c4-f82a7b6c2df9 · inbound

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis cites this paper.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:54.348620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:54.348620Z digest=sha256:796591768be088c1be523513604cdd7ec8a6ce57a90da60b7c29ab1934aabdb2

Observation 1fb147b1-feb1-4f5e-ab42-0d89a6116e3c · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T05:17:05.782526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:9c7c46af8b8e10d91934ce4bcf710da6010d81f92339e91b7501115aa1d6f9c6

Observation 2eceda81-b0f9-4057-8cee-6fd326be5e8f · inbound

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary cites this paper.

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:28.145339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:28.145339Z digest=sha256:4611a18cfa5823f14e448c0b8b3d9a1e5be7b4f4b115631a5fbf911369a84a63

Observation 1936df79-e57c-42bc-a82a-46f309452a11 · inbound

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities cites this paper.

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:35.194359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:35.194359Z digest=sha256:7d1cb1f2892224fba336019fa70df648d781d2d14941ac7832eb2ffd51defd46

Observation 9d88245e-cb38-489a-962f-89e09abc6976 · inbound

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning cites this paper.

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T23:25:45.324401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T23:24:43.556606Z digest=sha256:b1de84bb0475ec8b589ba36b63b5de0a38130d562193c02904082961206d7b71

Observation d2b52343-ad36-4d03-85ab-8d46ab45da4f · inbound

Towards Reliable, Uncertainty-Aware Alignment cites this paper.

Towards Reliable, Uncertainty-Aware Alignment Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:14.278132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:14.278132Z digest=sha256:2788d2b662537ddbb00d9e6cecf5675ab90a418ed0fe85635c8dc99d5e4eb28d

Observation d33daff4-48b0-4d19-ad27-e641f2b9cd80 · inbound

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training cites this paper.

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T22:40:43.264520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T22:38:57.833414Z digest=sha256:090a9dc08a912298c5d508c5137bdf6a91afcc9dc5f661cff3fca8b8dbcddcbb

Observation 2209bca8-a672-4f0c-963f-839ffae29ba4 · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T14:20:13.678202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:890909a47056bdf373de450c77aac970bc8516543d8f6ae32a8207870e1d16b0

Observation 53c188b7-ccb1-488a-a157-1e3c3af0e5bf · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.502437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.502437Z digest=sha256:bec8107db81c777211dd78b29134f88f913c6412790ce17ac8405b4d8de63269

Observation d1462243-f2f0-4201-a5d6-c14fc88c3639 · inbound

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models cites this paper.

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:33:16.862524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T20:29:31.354743Z digest=sha256:2156f9cccec0d6678d85ea27631dca1f97ce72ed118388f999eb7d08562adf0d

Observation 647e0a61-a9b6-445e-b119-5b1e41c6d209 · inbound

FUSE: Ensembling Verifiers with Zero Labeled Data cites this paper.

FUSE: Ensembling Verifiers with Zero Labeled Data Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:26:09.053345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T03:39:25.678190Z digest=sha256:11e9e84c9ad66c63d82618b431d47e4490f83664802cd1bcfc343c8fb368ccfb

Observation c2445045-e6d9-4182-b38a-59f9449287e6 · inbound

How Far Are Video Models from True Multimodal Reasoning? cites this paper.

How Far Are Video Models from True Multimodal Reasoning? Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:51:04.216352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:7e6f1613191952da9e128696492000c56ef9a4ee5c0dbe8bbe2311d261fbdc12

Observation 5dba58d5-f855-4cff-9966-6e019721a62f · inbound

Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring cites this paper.

Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T19:35:39.104230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T19:33:35.690030Z digest=sha256:91ec68967f91bc35be8cf530d3c089ecd6090bd5d92451483ce66ca073ed1230

Observation 91e20ee9-df19-4181-bc55-03a36ca23865 · inbound

Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring cites this paper.

Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:15:52.406421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:56:43.707252Z digest=sha256:dee744388fce20733f2a7d77466ee20a27e6452c0edbeec69b89d04447378bd5

Observation 595dd3df-fe53-4938-8421-b6fd18a0d474 · inbound

Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation cites this paper.

Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 126

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:56:08.381340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T16:53:00.860162Z digest=sha256:956c0a529152358d6e155c34b2a3edae57a3d08e71a7a6201facd4fecec9d15f

Observation b20806bf-3933-4d79-9015-e98990d3525c · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:45:59.518594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:1487d95d01a5f5f18aa24103efcde68194f850f5ed5953b0cb3555263390a793

Observation e39bdd7b-6892-43cb-a10e-85be05c618c2 · inbound

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? cites this paper.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:48:57.332773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:3986ce7e5097d43f846cf6c964259d46b8544c129f4b51988bd1722ebede6c20

Observation 8342c137-ccd7-40b3-9b84-bd3eec498c32 · inbound

The Human-AI Delegation-Verification Dilemma: Individual Strategies, Collective Equilibria and Sociotechnical Lock-in cites this paper.

The Human-AI Delegation-Verification Dilemma: Individual Strategies, Collective Equilibria and Sociotechnical Lock-in Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 17

Resolution
malformed identifier
arxiv_id, observed 2026-05-21T03:43:55.658050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T03:42:08.309581Z digest=sha256:f3b0ac35f32c289516fc7933d90b083098a4a931143c2aedb33bed479a1de618

Observation cb4d80df-6aaf-4f90-8445-5265e2833eb8 · inbound

The Human-AI Delegation-Verification Dilemma: Individual Strategies, Collective Equilibria and Sociotechnical Lock-in cites this paper.

The Human-AI Delegation-Verification Dilemma: Individual Strategies, Collective Equilibria and Sociotechnical Lock-in Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 17

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T17:14:56.535054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T17:07:54.749167Z digest=sha256:2c58793c380514a4edc605285ee1350ee1a742bbcf2cf3e8bb31c063fa4c6491

Observation c03e3297-ef02-4b98-b201-5ac4c2e9b141 · inbound

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs cites this paper.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:41:22.293334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T09:38:04.387777Z digest=sha256:ef7d08ae5816c8f73b5da7534e3a6fb12ded41a020ea65fbdc918071ca039cdd

Observation 683dab4d-0563-4100-970f-496a022b5b2b · inbound

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs cites this paper.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:41:21.380125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T09:38:04.387777Z digest=sha256:ce9f13250d86364d81a9d6a2419de7f363ceba16253146b8906b96bb1c65bffc

Observation 76842f9b-559e-43ae-8c8d-84afba033d7b · inbound

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs cites this paper.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:58.060030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:42c808e1556f178e31bcda38dc93c11c8a3337954866564c4a9d4134f1c606de

Observation 9ca641ca-ca54-47dc-9f8b-594dc499edfa · inbound

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models cites this paper.

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:46:26.792724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T11:32:16.166724Z digest=sha256:d1ea67d47dd171067639586dd47ea927721d35d4964063acd56c7f694a23abf5

Observation 16667bdb-4ecb-4f3a-816d-7d452f5c4284 · inbound

A Unifying Lens on Reward Uncertainty in RLHF cites this paper.

A Unifying Lens on Reward Uncertainty in RLHF Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:29.842084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T17:11:46.549150Z digest=sha256:e85509a3da0b80c0796d6ac9fd2a3e2f8677df637d3566e66617726d1a68d1d5

Observation 45b67a8f-0ef7-4aeb-9754-43663e65f0db · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.548957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:b52f451141492f5f926731531b330147cc65b18738710dd1efa91c0db36d11c8

Observation ce4e1c54-6b51-44d1-9195-3b9771a503d0 · inbound

Uncertainty-Aware Reward Modeling for Stable RLHF cites this paper.

Uncertainty-Aware Reward Modeling for Stable RLHF Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:30.697631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T18:14:33.673995Z digest=sha256:f747b65e4e5425db8a9eeb52b462464a12297114ab6c752cec152f6dd4533a2b

Observation 8b1a7393-77d3-4d5c-8dd4-545a86f3333d · inbound

Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation cites this paper.

Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:14:18.664913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:13:11.013395Z digest=sha256:7cb1b4e864864c2abb13bc8fcbf6446fa44bd3acc958b8655fecd14b96800490

Observation c52b93cc-4f9a-46e5-8677-0fc57a8e894f · inbound

Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation cites this paper.

Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:05:28.611341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T07:01:26.911110Z digest=sha256:124c72d10969c853a10be4c014102fb8054e91ae5e297f7089d3e2c57d5b278f

Observation 1589a814-1230-4d0d-a3e6-e6e0514900d8 · inbound

What do Reward Models Memorize? cites this paper.

What do Reward Models Memorize? Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:42.277767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:42.277767Z digest=sha256:e440013d20608a2a7ded7eb947a820b05756d33eaaf85fca5cfcf4c834e598d4