Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T18:37:18.229897Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 1 inbound Pith citation observation for arXiv:2605.04266.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T18:37:18.229897Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-13T02:39:02.891861Z
A source-named dated measurement, never combined with another source.
Source: cited_works
15 of 15 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a8981857-b564-41f4-a54d-d462b092b225 · outbound
Explaining and Preventing Alignment Collapse in Iterative RLHF P´ asztor et al
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7a66c6d1-6f9e-4504-8c9a-b6038792b215 · outbound
Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6255ff2c-3e68-433d-b930-72492ce0f157 · outbound
Explaining and Preventing Alignment Collapse in Iterative RLHF Figure 3 demonstrates that standard RLHF policies increasingly exploit noise dimensions, achieving high proxy rewards but abandoning true utility
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1508fd0d-ec0a-478c-b3b7-2e3297ea823e · outbound
Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5f0b6385-4c40-4b6c-9866-bc701ba4625a · outbound
Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5e054673-8b96-4f47-871e-45dc355466fa · outbound
Explaining and Preventing Alignment Collapse in Iterative RLHF The relaxed penalty multiplies this inner product by the overconfidence proxy σ(rϕ(x, yi) −r ϕ(x, y′)) −σ (U(x, yi) −U (x, y′)), where σ is the logistic sigmoid (cf
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 15032004-61a2-4ff3-be2b-5aa50e2b3eca · outbound
Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2aa457e6-936f-40b4-b569-bf4263836e0e · outbound
Explaining and Preventing Alignment Collapse in Iterative RLHF Both updates use the 8-bit paged AdamW optimizer [Dettmers et al., 2023, Loshchilov and Hutter, 2019] with gradients accumulated over four steps before each optimizer step
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 47ae9d58-1425-470e-85a4-d99848e67f82 · outbound
Explaining and Preventing Alignment Collapse in Iterative RLHF Punish models that confidently state false info
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5ccb1f38-95c8-4839-b8b2-4926e19552f9 · outbound
Explaining and Preventing Alignment Collapse in Iterative RLHF reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 492eaf3c-e480-44c5-9dc8-2d46961e4c97 · outbound
Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fa317903-27eb-4b84-9eb6-adaac08e533d · outbound
Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f8aa3efc-dc27-4243-8757-7331765b4e8e · outbound
Explaining and Preventing Alignment Collapse in Iterative RLHF spheri- cal
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation db703653-8228-47b4-8050-be4497ff7228 · outbound
Explaining and Preventing Alignment Collapse in Iterative RLHF Its training objective is the pointwise prediction error:ℓ train(z;ϕ) =ℓ(z;ϕ)
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 83ed3f0e-95ac-4523-bc44-85c0db65b96e · outbound
Explaining and Preventing Alignment Collapse in Iterative RLHF To frame reward maximization as a test loss to be minimized, we define the Leader’s objective as ℓtest(z; ϕ) = −rϕ(z)
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 81d5aaa8-a6b5-4379-9f87-0dd63c65de62 · inbound
Multimodal Reward Hacking in Reinforcement Learning Explaining and Preventing Alignment Collapse in Iterative RLHF
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.