Pith. sign in

Paper Citation Record · LEDGER

Confronting Reward Model Overoptimization with Constrained RLHF

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2310.04373.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.04373 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:56:03.454693Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:49:46.630731Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e15deb5d-b879-43a4-bc74-8f31f2df80e6 · inbound

Ctrl-DNA: Controllable Cell-Type-Specific Regulatory DNA Design via Constrained RL cites this paper.

Ctrl-DNA: Controllable Cell-Type-Specific Regulatory DNA Design via Constrained RL Confronting Reward Model Overoptimization with Constrained RLHF

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:03.454693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:03.454693Z digest=sha256:74e2a7050e64499ef9a0a1dce80156f89cf9516cf6f81fde143fee13c0041b41

Observation b6b24e2f-2f2f-40e5-8d64-81a0dae43c40 · inbound

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy cites this paper.

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy Confronting Reward Model Overoptimization with Constrained RLHF

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:30.335635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:30.335635Z digest=sha256:a7a0c81368afda225c1b17fa18e8d227dcfe6bc7a256502924dc8a289154df9d

Observation fbef0415-1371-4519-ac9b-cf890d9961a4 · inbound

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models cites this paper.

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models Confronting Reward Model Overoptimization with Constrained RLHF

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:58.050861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:58.050861Z digest=sha256:eab0f6b43030de5a9d8c5b97e854fb0c96c4cc264e1357d10974beb4059afc2a

Observation bae79e4c-a56b-4df0-a91f-d32ef34bc8cd · inbound

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models cites this paper.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Confronting Reward Model Overoptimization with Constrained RLHF

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.738852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.738852Z digest=sha256:a1103096cff2800c4d09a01de2638e690213d076820aad57d3e96e4f876de251

Observation e4b5172a-d780-42d1-9632-7ba3cc772345 · inbound

Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs cites this paper.

Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs Confronting Reward Model Overoptimization with Constrained RLHF

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:53:11.865838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T19:48:46.245576Z digest=sha256:1cc3491696cad17714fc9a188fa4d0a08f03a7939efe6de9f998aa78e616cdbd

Observation bc28d746-fc0a-4d3e-9bbd-631168517882 · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Confronting Reward Model Overoptimization with Constrained RLHF

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:25.597148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:412aa1e86437f3dd4be6f01fc8b93b5be967f1dca2250a8fd4e25b0068195426

Observation d2ce90b4-c8ce-42e6-887d-44f27f4e9163 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Confronting Reward Model Overoptimization with Constrained RLHF

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:57:06.410147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:6d9e84a8b094b1b4cf888b2b6dc752552f726a5f254e3b2ca70703ecf93126ec

Observation 816c00f9-199c-43b4-913b-a8a1d5f125f7 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Confronting Reward Model Overoptimization with Constrained RLHF

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:12:58.846443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:9cec383de2512899ba92fa24acb069a69231ffeda9f49bb11eb5ab16372c5286

Observation beb222e6-73b1-4e1f-8219-51073db9ad5a · inbound

Towards Context-Invariant Safety Alignment for Large Language Models cites this paper.

Towards Context-Invariant Safety Alignment for Large Language Models Confronting Reward Model Overoptimization with Constrained RLHF

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:39:40.886917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T05:36:12.562807Z digest=sha256:29fd378529b225484e2e003438c9becdf475e0ccfb5fd81ff78a4d64155ff409

Observation 25bac6a4-397c-4365-b5f1-d282b13ca5af · inbound

Against Proxy Optimization cites this paper.

Against Proxy Optimization Confronting Reward Model Overoptimization with Constrained RLHF

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:49:46.632296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T08:26:50.595370Z digest=sha256:5f58ab154a3e953a1492c45b32ca63d1159f1dc2d3a44db41b8c7fc95a6c9be9

Observation 3d5801d1-af89-46e9-9524-0791bd67cb62 · inbound

Spectral Rewiring for Exploration, Purification, and Model Merging cites this paper.

Spectral Rewiring for Exploration, Purification, and Model Merging Confronting Reward Model Overoptimization with Constrained RLHF

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T05:08:55.438431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:08:55.438431Z digest=sha256:af5679ce0fd71cc0d729cf3e9db9e4667af7a8ebf3206387f653e9e15f83aa66