Pith. sign in

Paper Citation Record · LEDGER

Explaining and Preventing Alignment Collapse in Iterative RLHF

As of 5 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 1 inbound Pith citation observation for arXiv:2605.04266.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.04266 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T18:37:18.229897Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T02:39:02.891861Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8981857-b564-41f4-a54d-d462b092b225 · outbound

This paper cites P´ asztor et al.

Explaining and Preventing Alignment Collapse in Iterative RLHF P´ asztor et al

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:27:23.559912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:1ad0e8f4e5bcd3dfc698d3718f793cc1a845459d5c381ad380d642c467b7ffd5

Observation 7a66c6d1-6f9e-4504-8c9a-b6038792b215 · outbound

This paper cites an unresolved cited work.

Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-26T04:27:23.571097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:03a686f82c83517bf7dbbcede121d574b253ca2da808cd009c69939db65ae92b

Observation 6255ff2c-3e68-433d-b930-72492ce0f157 · outbound

This paper cites Figure 3 demonstrates that standard RLHF policies increasingly exploit noise dimensions, achieving high proxy rewards but abandoning true utility.

Explaining and Preventing Alignment Collapse in Iterative RLHF Figure 3 demonstrates that standard RLHF policies increasingly exploit noise dimensions, achieving high proxy rewards but abandoning true utility

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:27:23.553283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:e409f56bf05967bc7259911083ac149a016ca50214c2e45b922bc8ac3504b2fa

Observation 1508fd0d-ec0a-478c-b3b7-2e3297ea823e · outbound

This paper cites an unresolved cited work.

Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-26T04:27:23.556353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:aa64cddd3b1481cd5dea1187c055d163ae4452041ee07f301e3bac256b2ca3e0

Observation 5f0b6385-4c40-4b6c-9866-bc701ba4625a · outbound

This paper cites an unresolved cited work.

Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-26T04:27:23.598294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:9ca4c85117613052ae03e19cc48253f3319abc8555ab73e2a96e5237ce438594

Observation 5e054673-8b96-4f47-871e-45dc355466fa · outbound

This paper cites The relaxed penalty multiplies this inner product by the overconfidence proxy σ(rϕ(x, yi) −r ϕ(x, y′)) −σ (U(x, yi) −U (x, y′)), where σ is the logistic sigmoid (cf.

Explaining and Preventing Alignment Collapse in Iterative RLHF The relaxed penalty multiplies this inner product by the overconfidence proxy σ(rϕ(x, yi) −r ϕ(x, y′)) −σ (U(x, yi) −U (x, y′)), where σ is the logistic sigmoid (cf

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:27:23.584703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:34ab9582c578577b7862fc3e4012677d4ac913c5dbce7a7744e1e82abbc675c3

Observation 15032004-61a2-4ff3-be2b-5aa50e2b3eca · outbound

This paper cites an unresolved cited work.

Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-26T04:27:23.594988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:b25d751f799793248f466a1bbf099b330f56f16daac4f074bf068d32698450a5

Observation 2aa457e6-936f-40b4-b569-bf4263836e0e · outbound

This paper cites Both updates use the 8-bit paged AdamW optimizer [Dettmers et al., 2023, Loshchilov and Hutter, 2019] with gradients accumulated over four steps before each optimizer step.

Explaining and Preventing Alignment Collapse in Iterative RLHF Both updates use the 8-bit paged AdamW optimizer [Dettmers et al., 2023, Loshchilov and Hutter, 2019] with gradients accumulated over four steps before each optimizer step

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:27:23.548930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:28649dd7a39f5d77d94065aa864a33b204e29342a18306e12aee003b733c95e6

Observation 47ae9d58-1425-470e-85a4-d99848e67f82 · outbound

This paper cites Punish models that confidently state false info.

Explaining and Preventing Alignment Collapse in Iterative RLHF Punish models that confidently state false info

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:27:23.588269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:762f88f33f9a52a0897ef1fd2869326fb967fbeea6a75507366e40e48dc901ad

Observation 5ccb1f38-95c8-4839-b8b2-4926e19552f9 · outbound

This paper cites reasoning.

Explaining and Preventing Alignment Collapse in Iterative RLHF reasoning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:27:23.577637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:dd90bcbc955811dedcded367758c4eda9be99895aef5c5662b703183d1d6270a

Observation 492eaf3c-e480-44c5-9dc8-2d46961e4c97 · outbound

This paper cites an unresolved cited work.

Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-26T04:27:23.591621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:560799c5fda9f5e2176101e1b44e1c226269ab3e020462e6b6ce49a080fd1411

Observation fa317903-27eb-4b84-9eb6-adaac08e533d · outbound

This paper cites an unresolved cited work.

Explaining and Preventing Alignment Collapse in Iterative RLHF Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-26T04:27:23.574459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:f960bcbd215bc3444a11b0d4a1ea66ec3ad09569233d9935d591afca0091335f

Observation f8aa3efc-dc27-4243-8757-7331765b4e8e · outbound

This paper cites spheri- cal.

Explaining and Preventing Alignment Collapse in Iterative RLHF spheri- cal

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:27:23.581112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:e7c55cdf817550a0fc343ae71fba36e0ae1950960cc82132f3405a96e95e4537

Observation db703653-8228-47b4-8050-be4497ff7228 · outbound

This paper cites Its training objective is the pointwise prediction error:ℓ train(z;ϕ) =ℓ(z;ϕ).

Explaining and Preventing Alignment Collapse in Iterative RLHF Its training objective is the pointwise prediction error:ℓ train(z;ϕ) =ℓ(z;ϕ)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:27:23.567753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:ed507adcea236fc5f4ddcc407f58f48bec367f42f483dd95e869c16a5eb88dee

Observation 83ed3f0e-95ac-4523-bc44-85c0db65b96e · outbound

This paper cites To frame reward maximization as a test loss to be minimized, we define the Leader’s objective as ℓtest(z; ϕ) = −rϕ(z).

Explaining and Preventing Alignment Collapse in Iterative RLHF To frame reward maximization as a test loss to be minimized, we define the Leader’s objective as ℓtest(z; ϕ) = −rϕ(z)

Reference 15

Resolution
malformed identifier
raw_fallback, observed 2026-05-26T04:27:23.563431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:37:18.229897Z digest=sha256:20ce792edc3df88034c89f637718f8a5583733ecd8c80058700e1f5439d061b8

Pith citing papers

Observation 81d5aaa8-a6b5-4379-9f87-0dd63c65de62 · inbound

Multimodal Reward Hacking in Reinforcement Learning cites this paper.

Multimodal Reward Hacking in Reinforcement Learning Explaining and Preventing Alignment Collapse in Iterative RLHF

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:d9fc3b4b4d93ece4f17746d621d2467fedbb775c40f564553189b77af5b49d26