Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2310.04373.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:56:03.454693Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:49:46.630731Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation e15deb5d-b879-43a4-bc74-8f31f2df80e6 · inbound
Ctrl-DNA: Controllable Cell-Type-Specific Regulatory DNA Design via Constrained RL Confronting Reward Model Overoptimization with Constrained RLHF
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6b24e2f-2f2f-40e5-8d64-81a0dae43c40 · inbound
Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy Confronting Reward Model Overoptimization with Constrained RLHF
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbef0415-1371-4519-ac9b-cf890d9961a4 · inbound
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models Confronting Reward Model Overoptimization with Constrained RLHF
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bae79e4c-a56b-4df0-a91f-d32ef34bc8cd · inbound
MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Confronting Reward Model Overoptimization with Constrained RLHF
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4b5172a-d780-42d1-9632-7ba3cc772345 · inbound
Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs Confronting Reward Model Overoptimization with Constrained RLHF
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc28d746-fc0a-4d3e-9bbd-631168517882 · inbound
Reinforcement Learning via Value Gradient Flow Confronting Reward Model Overoptimization with Constrained RLHF
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d2ce90b4-c8ce-42e6-887d-44f27f4e9163 · inbound
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Confronting Reward Model Overoptimization with Constrained RLHF
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 816c00f9-199c-43b4-913b-a8a1d5f125f7 · inbound
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Confronting Reward Model Overoptimization with Constrained RLHF
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation beb222e6-73b1-4e1f-8219-51073db9ad5a · inbound
Towards Context-Invariant Safety Alignment for Large Language Models Confronting Reward Model Overoptimization with Constrained RLHF
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25bac6a4-397c-4365-b5f1-d282b13ca5af · inbound
Against Proxy Optimization Confronting Reward Model Overoptimization with Constrained RLHF
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d5801d1-af89-46e9-9524-0791bd67cb62 · inbound
Spectral Rewiring for Exploration, Purification, and Model Merging Confronting Reward Model Overoptimization with Constrained RLHF
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.