Pith. sign in

Paper Citation Record · LEDGER

Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2405.16436.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.16436 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:44:57.905368Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T15:51:29.355251Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5370dfe7-0ac7-412f-ac8c-257b3516106c · inbound

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs cites this paper.

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 267

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:51:29.358597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T15:51:29.022336Z digest=sha256:186d7c62554ee9f93388c1a7cd61e0ddd739bfd2a9405fcd6ce58a8785df69fe

Observation 206d38f8-6001-4688-a32f-53524120eb29 · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.905368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.905368Z digest=sha256:54a6f89d75e195da7fe4f9ef31eb379233f92045cbdfc24fb3bfca766007381f

Observation 551e7d34-90c3-4fd9-adf5-b9717b571467 · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:47.390690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:47.390690Z digest=sha256:6239e46bd6749c1bbdf9f7566eadb0bbc6139005f5af04be71eccc6a32d69849

Observation ae148c67-3aee-4709-bed9-cfe1c368eacc · inbound

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment cites this paper.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:38.224309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:38.224309Z digest=sha256:13a85f1eb2160e4a7bfcef7459532d618453f54b6b5d783cc33604af0afad773

Observation 4b856fd7-88d8-4400-897b-592296f9adef · inbound

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion cites this paper.

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T04:36:13.529125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:36:13.529125Z digest=sha256:8b44d4ec37253e1abed4532bb5d757a1f4d59d96cbcc11bb0e1c2f215088028a

Observation 72162c4c-fb81-48bd-96d3-f3e60f6772bc · inbound

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models cites this paper.

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.658973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:35:13.659698Z digest=sha256:b52e257563059fcb9b5d7bcf08925e1ce05429366d36de5b0b4801d58e38a8bd

Observation 1a6b039b-9246-4efc-8272-be4ed36d7e4b · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:55439b63c22f22e78b4aff2a91d2b1c9bf1570d4e1915f3c0df54b73cd711622

Observation 779f7c7f-44bf-413e-bb64-6a5a6ac6338b · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:41.047599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:41.047599Z digest=sha256:424564f03f419d2dfa577a2d8672d547e6e00ef7966f55054ac04cf248f763f3