Pith. sign in

Paper Citation Record · LEDGER

Reward Model Ensembles Help Mitigate Overoptimization

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2310.02743.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.02743 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:57:54.254039Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 56747fdd-9082-453d-8616-4d7b19198df2 · inbound

InfAlign: Inference-aware language model alignment cites this paper.

InfAlign: Inference-aware language model alignment Reward Model Ensembles Help Mitigate Overoptimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:57:54.254039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:57:54.254039Z digest=sha256:f8143fed312163a0a59d123d08dee3f2c2def6cb34960078d4db00aea16e84ed

Observation 3404ba26-329b-4d9c-837a-3c7f7aaaffc3 · inbound

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment cites this paper.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Reward Model Ensembles Help Mitigate Overoptimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.559707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.559707Z digest=sha256:4124a29619e5d6023e16ec6ae1298f0796848933f2f41f15bcc14e323de6bb71

Observation a093c96a-6bff-4f54-83cc-876ca28f9af7 · inbound

Debate Helps Weak-to-Strong Generalization cites this paper.

Debate Helps Weak-to-Strong Generalization Reward Model Ensembles Help Mitigate Overoptimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T17:50:56.339928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:50:56.339928Z digest=sha256:e610727b290ee3c88d12486407f176f1b2c2d778136ceb2e164fc7c2c49dc9e7

Observation e538dc2d-97f4-4c23-8dd9-ab195e63b329 · inbound

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment cites this paper.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Reward Model Ensembles Help Mitigate Overoptimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.522690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.522690Z digest=sha256:08bfbe2f8e455f6d8b2e89a70ea46b01e116d935b887cb0be00152cf98a95f70

Observation 5258d64d-76db-48d4-accb-5a59a590a8eb · inbound

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs cites this paper.

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Reward Model Ensembles Help Mitigate Overoptimization

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-09T11:32:47.837706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:32:47.837706Z digest=sha256:944803e30aafb76c9355081d5e19a40297d9c1fb97deb5a483696c41dcd4cf14

Observation 0c3fe5c1-4a8a-434b-865d-c3999a37b8c8 · inbound

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning cites this paper.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Reward Model Ensembles Help Mitigate Overoptimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.051139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.051139Z digest=sha256:96be2e3b381db90a765144310b6e252bf1fefc0a76d115b8c8df9c3b4b0b09fd

Observation 726b7ead-80f1-4df5-a9f2-480b78a979cd · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models Reward Model Ensembles Help Mitigate Overoptimization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.902238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.902238Z digest=sha256:b2dd17bacd602039374e8ba364fba03cbe69050ef7a4cab0b5070bff5dd091af

Observation 1560d3f1-fc0a-4acf-8370-8e9b79abaaec · inbound

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary cites this paper.

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Reward Model Ensembles Help Mitigate Overoptimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:30.200216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:30.200216Z digest=sha256:0044c39450e90e2349da297af1e2a69b9787c22343c3f78577198bea3c60d828

Observation 12d4acc6-c2fe-47e3-8ed9-3c0a9e4d5f3e · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Reward Model Ensembles Help Mitigate Overoptimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:24.968234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:24.968234Z digest=sha256:ce46295ab7960035d043a5883c04ad5c443dbcf8e3ef68db024278d356fe13a4

Observation d81661bf-1992-4d60-91f9-a40889259d1e · inbound

Towards Reliable, Uncertainty-Aware Alignment cites this paper.

Towards Reliable, Uncertainty-Aware Alignment Reward Model Ensembles Help Mitigate Overoptimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:13.721341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:13.721341Z digest=sha256:0caddad8e292eb8d14884ad003ab5e86377f7553e9ccaabf061ed04e7935f5f0

Observation 32ede528-99af-4867-8c7c-71b624daa0fc · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF Reward Model Ensembles Help Mitigate Overoptimization

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:20:13.627586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:db00550ca08950f5bc46043f60dd94f548d484c7ae7026c8ed9495e8ef47d91e

Observation b7eced65-43f1-4476-b2e8-5780b93e53dc · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Reward Model Ensembles Help Mitigate Overoptimization

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.509853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:75de775715a3b56e74b013d219ee88b5d54a8d991b59c49103ad4580df49220b

Observation 7146a6c9-53c9-47d7-a33f-47bb6137bb49 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Reward Model Ensembles Help Mitigate Overoptimization

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:10:10.523375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:c8615f7f677c2752836c0a476f7854a81054dd78e775806d5a4341a8eba9f3fb

Observation 96982241-62cb-40e8-b1b3-331ca867a0ba · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Reward Model Ensembles Help Mitigate Overoptimization

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:25.590488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:5c21c81ef4f2a76b0540b3a5ac767d84715877efdc5b6a5d21374154b14a69f1

Observation 4c002d6f-dab3-463b-8179-4bf45742812e · inbound

FUSE: Ensembling Verifiers with Zero Labeled Data cites this paper.

FUSE: Ensembling Verifiers with Zero Labeled Data Reward Model Ensembles Help Mitigate Overoptimization

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:26:09.036855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T03:39:25.678190Z digest=sha256:9f06a36eb50cd7bd234c724a58a46e246c04d39f77e6fc3ae18a4e8c34b8064a

Observation 21372565-3424-40f8-a0fe-139255f3463c · inbound

Theoretical Limits of Language Model Alignment cites this paper.

Theoretical Limits of Language Model Alignment Reward Model Ensembles Help Mitigate Overoptimization

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:56.966652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T01:18:37.614335Z digest=sha256:b57a3637e8db3eec4e2dfde88eab8a1a0a84eef4f534d4b9e83f4c2132689a52

Observation 48120300-4e5a-45c6-b2a4-d5056048f549 · inbound

Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants cites this paper.

Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants Reward Model Ensembles Help Mitigate Overoptimization

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:36:46.897334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T02:28:13.317630Z digest=sha256:7bc61dc913dc70e19bdc17a4fdd651d8936b88606fd2c37d057a50c662735c10

Observation 5ed0fe47-9729-45d5-b6c2-bebe997b37a6 · inbound

TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment cites this paper.

TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment Reward Model Ensembles Help Mitigate Overoptimization

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:37:29.192797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T07:36:16.811765Z digest=sha256:ef82db686e0f0887f62315012b851832cb7df6f4e338a88795223871036f1d5c

Observation 44d10863-751c-4a77-bdd7-f20068c52e47 · inbound

TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment cites this paper.

TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment Reward Model Ensembles Help Mitigate Overoptimization

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:03:02.908788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T22:01:21.695270Z digest=sha256:3626b480fbf4c56684e14bb1a0f0c058a05e834602f9bff6f727def55563e805

Observation 6e2d1bdb-71b1-4184-acae-d07b2895a53b · inbound

Variance-aware Reward Modeling with Anchor Guidance cites this paper.

Variance-aware Reward Modeling with Anchor Guidance Reward Model Ensembles Help Mitigate Overoptimization

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:12:17.505258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-13T05:11:05.546092Z digest=sha256:b667bc35767d26b95e6e11257fca32278e7785ad1d6e651f8fba9db37419bc17

Observation cbef8efc-0e05-4646-a69f-dc02226ae345 · inbound

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs cites this paper.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Reward Model Ensembles Help Mitigate Overoptimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:41:22.253533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-22T09:38:04.387777Z digest=sha256:6d743a9851a26d164c61158bce8ad1110f14742e46fb6bba205a0d864efde106

Observation f2ebf4ad-b424-4b6e-a21d-1fed4734bd6e · inbound

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs cites this paper.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Reward Model Ensembles Help Mitigate Overoptimization

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:41:21.386002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-22T09:38:04.387777Z digest=sha256:320c1b331bf381d2e8fcbc7a85c5c8e702cc006bc6d03f01880353a19a9e9c93

Observation c0c514f6-1221-492d-82a9-150c6bb749d0 · inbound

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs cites this paper.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Reward Model Ensembles Help Mitigate Overoptimization

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:58.057414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:a84b4270417038ef99b3ea2271aa74f8d871875917fddcc3a73a08d12a718c11

Observation 926e9c09-7e2b-4af4-b7dc-b6714b222c07 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Reward Model Ensembles Help Mitigate Overoptimization

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.389148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:5499690a39b9ac610a4cd2c1710ca1757d9714b02cba7bf0e39c4f70bea2c847

Observation ef72fedb-2c99-414b-baf1-f5c5b96e8386 · inbound

Uncertainty-Aware Reward Modeling for Stable RLHF cites this paper.

Uncertainty-Aware Reward Modeling for Stable RLHF Reward Model Ensembles Help Mitigate Overoptimization

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:19:30.702713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T18:14:33.673995Z digest=sha256:e01c6cee0eeea52364030bc9b4628d413fe8537b6206a630001170baa2377a1d

Observation d6cad619-e42a-43b6-bd45-7ad7a0fa5bd8 · inbound

Against Proxy Optimization cites this paper.

Against Proxy Optimization Reward Model Ensembles Help Mitigate Overoptimization

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:49:46.643173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T08:26:50.595370Z digest=sha256:4422a0293d0aeb712a0de366153563d72ccb2e2b3073131b00d30333cabc0b1f

Observation d9ac9227-cc96-4ce7-8e35-5affe4476085 · inbound

Designing Reward Signals for Portable Query Generation: A Case Study in Industrial Semantic Job Search cites this paper.

Designing Reward Signals for Portable Query Generation: A Case Study in Industrial Semantic Job Search Reward Model Ensembles Help Mitigate Overoptimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:29:51.705362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T05:10:59.825246Z digest=sha256:4463e42ca5e7857c169f0e888d70a8bf3d752f7dfb5522cd194de4a2f51c19fd

Observation 325d1168-d703-44be-b697-bc9e1cf32054 · inbound

STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning cites this paper.

STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning Reward Model Ensembles Help Mitigate Overoptimization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:24:19.479860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T06:17:48.731091Z digest=sha256:2cdacf3e581738220a8505b098344f51252596775d0b5e0fe4f31633cdedb81e

Observation 50ba71a1-39ee-4081-94da-de8cc3d99fab · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Reward Model Ensembles Help Mitigate Overoptimization

Reference 283

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:0cb3895d0df3321f9c00a6387dad45759e08ad038de045c29c4943e6560ec9db

Observation d41dcab2-c667-4af5-a4dc-773fd6247303 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Reward Model Ensembles Help Mitigate Overoptimization

Reference 284

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:05.531058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:05.531058Z digest=sha256:cff6bc66ef255f91ae9212f2b2d77449949725becb47e6a6ce526d4dfcc7f6eb

Observation e426a20c-f573-4c4b-a2db-4c6fccabcff7 · inbound

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges cites this paper.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Reward Model Ensembles Help Mitigate Overoptimization

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T21:25:38.636176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:7fde7c7674e179992e213495dac81993c2f2085ba89c491378f36df4315793aa