Pith. sign in

Paper Citation Record · LEDGER

Confronting Reward Model Overoptimization with Constrained RLHF

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2310.04373.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.04373 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:56:03.454693Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:49:46.630731Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e15deb5d-b879-43a4-bc74-8f31f2df80e6 · inbound

Ctrl-DNA: Controllable Cell-Type-Specific Regulatory DNA Design via Constrained RL cites this paper.

Ctrl-DNA: Controllable Cell-Type-Specific Regulatory DNA Design via Constrained RL Confronting Reward Model Overoptimization with Constrained RLHF

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:03.454693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:03.454693Z digest=sha256:b10e877b308e2c921a301eaa80f0ca709efe93b27c5c761ed6493a9217fe537a

Observation b6b24e2f-2f2f-40e5-8d64-81a0dae43c40 · inbound

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy cites this paper.

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy Confronting Reward Model Overoptimization with Constrained RLHF

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:30.335635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:30.335635Z digest=sha256:bb4c8dda14caaa3e84b92653eab74d976356cee4e92ed8dbe4695e64baeda1fe

Observation fbef0415-1371-4519-ac9b-cf890d9961a4 · inbound

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models cites this paper.

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models Confronting Reward Model Overoptimization with Constrained RLHF

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:58.050861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:58.050861Z digest=sha256:eab0f6b43030de5a9d8c5b97e854fb0c96c4cc264e1357d10974beb4059afc2a

Observation bae79e4c-a56b-4df0-a91f-d32ef34bc8cd · inbound

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models cites this paper.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Confronting Reward Model Overoptimization with Constrained RLHF

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.738852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.738852Z digest=sha256:9386a4f43826c9676a4da778bce0d1c08b1806b4dbd935612faaf8a66c46c035

Observation e4b5172a-d780-42d1-9632-7ba3cc772345 · inbound

Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs cites this paper.

Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs Confronting Reward Model Overoptimization with Constrained RLHF

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:53:11.865838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T19:48:46.245576Z digest=sha256:94149a93482dbc40508531822de90080227ceb893674944ca06663a43e1a3215

Observation bc28d746-fc0a-4d3e-9bbd-631168517882 · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Confronting Reward Model Overoptimization with Constrained RLHF

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:25.597148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:a103e64b35ffa6d403ed0bcab9a284be765dfb223913ef3586558f735861f071

Observation d2ce90b4-c8ce-42e6-887d-44f27f4e9163 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Confronting Reward Model Overoptimization with Constrained RLHF

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:57:06.410147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:15384b6c1860ec12412e4d13901a8ac5f7179d83b84e302d3a38a33c860a24d4

Observation 816c00f9-199c-43b4-913b-a8a1d5f125f7 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Confronting Reward Model Overoptimization with Constrained RLHF

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:12:58.846443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:04c665fb4e9c9876814a42ee3018a90bb945bbd0eaaa252dc404de48214f5bd9

Observation beb222e6-73b1-4e1f-8219-51073db9ad5a · inbound

Towards Context-Invariant Safety Alignment for Large Language Models cites this paper.

Towards Context-Invariant Safety Alignment for Large Language Models Confronting Reward Model Overoptimization with Constrained RLHF

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:39:40.886917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T05:36:12.562807Z digest=sha256:bd7162818203dbf94fa4c4bbc209d4a6615fee99399bd9124201283b70188adf

Observation 25bac6a4-397c-4365-b5f1-d282b13ca5af · inbound

Against Proxy Optimization cites this paper.

Against Proxy Optimization Confronting Reward Model Overoptimization with Constrained RLHF

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:49:46.632296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T08:26:50.595370Z digest=sha256:cf672a0ecc00ded43ee984282e876571d96d9a2f7511bbea0455f4c63ea9e118

Observation 3d5801d1-af89-46e9-9524-0791bd67cb62 · inbound

Spectral Rewiring for Exploration, Purification, and Model Merging cites this paper.

Spectral Rewiring for Exploration, Purification, and Model Merging Confronting Reward Model Overoptimization with Constrained RLHF

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T05:08:55.438431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:08:55.438431Z digest=sha256:1c36ac9c5a0a64767fb3017d826b7ffaee910e3b9ff62d318f03ade409ad347e