Pith. sign in

Paper Citation Record · LEDGER

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling

As of 8 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 1 inbound Pith citation observation for arXiv:2506.08681.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08681 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:16:43.217090Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-25T03:41:52.859647Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T03:45:17.571390Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a0f685e8-a465-468b-a58f-4832787db942 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Scaling Laws for Reward Model Overoptimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.181106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.181106Z digest=sha256:648dddee7eed594d46ec3a6c10fd12c0f7488c755b9203119cbb1caa224de76b

Observation 258e424f-f7ff-40c9-a1f4-26cfa3be0f3d · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Direct Language Model Alignment from Online AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.185377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.185377Z digest=sha256:570b02c585be8290379d5da7e7334911419f9c9e3fd862eaba1a2959db7b7928

Observation 4d66ed61-f61e-490b-ae51-adb4cca2b999 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Direct Language Model Alignment from Online AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.189540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.189540Z digest=sha256:30e946ecfd0368a75315df3b291cd491034601c442f5e2dd9d4668f92f11b844

Observation 2ecd451f-e590-4233-8ca6-48f70ef5035f · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.197365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.197365Z digest=sha256:a24ec9d953176cc3bffce382f49d23bb2a05376bba6d3858907768230ec6b630

Observation f2f2f166-723e-4ae7-9a6c-85d85b5b8c91 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.378995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:16:43.201513Z digest=sha256:6ca2a7b16bf1b77bc02dba7fe005e890c918d65f7db5e6f5d0072ce0f600e6b2

Observation 847a3684-3bc7-45c9-9457-2fc086b44a42 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Proximal Policy Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.209489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.209489Z digest=sha256:ee75e6d433fb88a6b23b3e7d68a954101d642b2849212d1d245c6fe6b682451e

Observation b43310e1-adb9-47f2-ba3b-74fa4eb505fe · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.353429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:16:43.213207Z digest=sha256:d3be46eda6dd40404a1afb81ae1eb9c4655a2d9273853794630c545c0ad39f9e

Observation 9c79d974-d0d4-4636-9e46-2afef4379b97 · outbound

This paper cites gpt-4o-mini.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling gpt-4o-mini

Reference 640

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:16:43.340209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:16:43.217090Z digest=sha256:44853317aab24d33ddee08efec1dddf6e9f7086bc26ff891f6bed6f2d11ef9a4

Observation 6c2a0880-c598-4a3f-b7f5-1f9f2cf431ed · outbound

This paper cites cc/paper_files/paper/2017/file/d5e2c0adad503c91f91df240d0cd4e49-Paper.pdf.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling cc/paper_files/paper/2017/file/d5e2c0adad503c91f91df240d0cd4e49-Paper.pdf

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:16:43.405220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:16:43.171719Z digest=sha256:10f35f2a0242b69aded25bb75f1236b95a34fb2d0d945648a8509a24f848df4f

Observation a77382f5-edcc-4ce8-af9b-2e9f9ac1bef2 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.391070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:16:43.193528Z digest=sha256:9ca354e95ca2f29b15d9e2ba31fc8fd289c1cc41fa928518dc86c5d6ec50f31a

Observation 220cbbf9-df1c-41a9-9a11-0e648098282f · outbound

This paper cites doi.org/10.1002/9781118445112.stat08284.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling doi.org/10.1002/9781118445112.stat08284

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.176803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.176803Z digest=sha256:05c6f3eb4cc19966d9282d3c22446d76f7f3b0a7992a8951eb92160e3a3cd0f6

Observation a041404c-0989-4b13-b3a2-30e20740b42a · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.162886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.162886Z digest=sha256:bdabed7b1d97dc8fddd7d17c3af5d646160b20bec3953608b561f905be8d96d8

Observation 4a290447-6663-444f-9ffa-cd2c74a00af0 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.365925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:16:43.205323Z digest=sha256:facb3aefb9300659c4d5600a0fd56c9dbbd9755c1cee1d6d0ca9c13a466d0883

Observation e92b9e93-09fb-4fc7-a05a-db9669aceec6 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.157955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.157955Z digest=sha256:949d016704d1ac8d98693a146744f05b88a90688f338b3c91eaf9754290ab10e

Observation d41e3211-0888-4820-835d-ee00b26f148a · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.418032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:16:43.167292Z digest=sha256:c64f6796bb126c1ce076d571febad0e6f20e3b0df2e27c47b4f5e5a8242c5f1f

Pith citing papers

Observation e1f65f80-3589-45c1-a711-d6626b6e7187 · inbound

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization cites this paper.

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:45:17.573771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T03:41:52.859647Z digest=sha256:f3de7166b79747d0be735b998ba6e5aaa337bd7354b814fa33043c0d66fe3782