Pith. sign in

Paper Citation Record · LEDGER

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling

As of 8 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 1 inbound Pith citation observation for arXiv:2506.08681.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08681 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:16:43.217090Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-25T03:41:52.859647Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T03:45:17.571390Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a0f685e8-a465-468b-a58f-4832787db942 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Scaling Laws for Reward Model Overoptimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.181106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.181106Z digest=sha256:648dddee7eed594d46ec3a6c10fd12c0f7488c755b9203119cbb1caa224de76b

Observation 258e424f-f7ff-40c9-a1f4-26cfa3be0f3d · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Direct Language Model Alignment from Online AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.185377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.185377Z digest=sha256:570b02c585be8290379d5da7e7334911419f9c9e3fd862eaba1a2959db7b7928

Observation 4d66ed61-f61e-490b-ae51-adb4cca2b999 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Direct Language Model Alignment from Online AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.189540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.189540Z digest=sha256:30e946ecfd0368a75315df3b291cd491034601c442f5e2dd9d4668f92f11b844

Observation 2ecd451f-e590-4233-8ca6-48f70ef5035f · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.197365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.197365Z digest=sha256:a24ec9d953176cc3bffce382f49d23bb2a05376bba6d3858907768230ec6b630

Observation f2f2f166-723e-4ae7-9a6c-85d85b5b8c91 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.378995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:16:43.201513Z digest=sha256:d8fb9c7efbfe305ff273f43c14e7eed5e87143334cef415f1eb7632cdbef5675

Observation 847a3684-3bc7-45c9-9457-2fc086b44a42 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Proximal Policy Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.209489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.209489Z digest=sha256:ee75e6d433fb88a6b23b3e7d68a954101d642b2849212d1d245c6fe6b682451e

Observation b43310e1-adb9-47f2-ba3b-74fa4eb505fe · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.353429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:16:43.213207Z digest=sha256:891ce8dcf677c1d31451b03c8ebb416d7991de81fd6715cac2bbe453a5e0ac39

Observation 9c79d974-d0d4-4636-9e46-2afef4379b97 · outbound

This paper cites gpt-4o-mini.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling gpt-4o-mini

Reference 640

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:16:43.340209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:16:43.217090Z digest=sha256:7022162ea6e303b5ab976f3cda31127e30d57009416fec1546478e8da2a1f970

Observation 6c2a0880-c598-4a3f-b7f5-1f9f2cf431ed · outbound

This paper cites cc/paper_files/paper/2017/file/d5e2c0adad503c91f91df240d0cd4e49-Paper.pdf.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling cc/paper_files/paper/2017/file/d5e2c0adad503c91f91df240d0cd4e49-Paper.pdf

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:16:43.405220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:16:43.171719Z digest=sha256:287d8f6ffd3a1fd6bc4c4e9cd2f4c95bedebae96bda895db5ab738966a8d9497

Observation a77382f5-edcc-4ce8-af9b-2e9f9ac1bef2 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.391070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:16:43.193528Z digest=sha256:32c017ced82ef7412cdeb8432eca4a8a5ed1b01b338b5dc1bcbdc1f60dbaa807

Observation 220cbbf9-df1c-41a9-9a11-0e648098282f · outbound

This paper cites doi.org/10.1002/9781118445112.stat08284.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling doi.org/10.1002/9781118445112.stat08284

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.176803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.176803Z digest=sha256:05c6f3eb4cc19966d9282d3c22446d76f7f3b0a7992a8951eb92160e3a3cd0f6

Observation a041404c-0989-4b13-b3a2-30e20740b42a · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.162886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.162886Z digest=sha256:bdabed7b1d97dc8fddd7d17c3af5d646160b20bec3953608b561f905be8d96d8

Observation 4a290447-6663-444f-9ffa-cd2c74a00af0 · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.365925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:16:43.205323Z digest=sha256:d249a070ca8207b6cba440b56806b2d3c0d65576c0b52d5e6894e22c218f9bd0

Observation e92b9e93-09fb-4fc7-a05a-db9669aceec6 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:16:43.157955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:16:43.157955Z digest=sha256:949d016704d1ac8d98693a146744f05b88a90688f338b3c91eaf9754290ab10e

Observation d41e3211-0888-4820-835d-ee00b26f148a · outbound

This paper cites an unresolved cited work.

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:16:43.418032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:16:43.167292Z digest=sha256:2d4b13ac16f8d0b34f6f38389352e3b84dfb409886d4886a172312c0ad9cb5b5

Pith citing papers

Observation e1f65f80-3589-45c1-a711-d6626b6e7187 · inbound

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization cites this paper.

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:45:17.573771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T03:41:52.859647Z digest=sha256:010dbcfdef93cca00aea87a6d489613274730198eea730fd77ad6538722e04a7