Pith. sign in

Paper Citation Record · LEDGER

Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2406.02900.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.02900 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:43.718310Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T21:25:38.650008Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3ed7ad6e-b5db-4026-8e0b-6a6ffabe27d5 · inbound

Efficient Alignment of Large Language Models via Data Sampling cites this paper.

Efficient Alignment of Large Language Models via Data Sampling Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T19:40:15.662140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:40:15.662140Z digest=sha256:d07a7dcd273d4240f14e93676ec81f5f81200de6d5f6bdc7ffa3b1a61a57e385

Observation 6113cc10-fc17-433c-ba21-e4283a48f874 · inbound

Does RLHF Scale? Exploring the Impacts From Data, Model, and Method cites this paper.

Does RLHF Scale? Exploring the Impacts From Data, Model, and Method Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T20:10:16.120361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:10:16.120361Z digest=sha256:bafd17c607f9c3e0ab1b4c9e8c0bba71ea3c3449a00f3bb16d86fa3fa0c02a64

Observation 7c46f08b-c16d-4d5b-b3dc-097a730d8a2d · inbound

When Can Proxies Improve the Sample Complexity of Preference Learning? cites this paper.

When Can Proxies Improve the Sample Complexity of Preference Learning? Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.040308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.040308Z digest=sha256:21c131bcc374ead01f2b32be2f0049ffe7d1dcbcaa0b3df9703b6a6ad1b47ff9

Observation b9ef3da2-eead-474a-8b86-200a858c4dc4 · inbound

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization cites this paper.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.662433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.662433Z digest=sha256:38627cd23d28f8d38b632c70c4ebb7baa3212aefc306a34dcafbe1b67dffdf3a

Observation ea404f9e-bb08-4250-989e-7ef428c33648 · inbound

AlphaPO: Reward Shape Matters for LLM Alignment cites this paper.

AlphaPO: Reward Shape Matters for LLM Alignment Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T21:51:08.647086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:51:08.647086Z digest=sha256:df5cff06f3560afafe2475ded27c218191dc3572329a2d2fba287cae9122c05b

Observation fb2d9cc5-83f4-4ffb-a3af-d78b90e26c4c · inbound

T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling cites this paper.

T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T18:05:43.122286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:05:43.122286Z digest=sha256:f0d46336fbc77c7adb81a777ff51d6a720e0b8692caf59099864b828c899eddd

Observation dc65dccf-c32c-4e79-840d-5f81612232a5 · inbound

Online Preference Alignment for Language Models via Count-based Exploration cites this paper.

Online Preference Alignment for Language Models via Count-based Exploration Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.560951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.560951Z digest=sha256:b68441abe52b3e5099e0cd839001f564f07becdf484329e0af0608bcea406654

Observation aa73abb2-54c2-4e48-a89f-a33772d9189b · inbound

Debate Helps Weak-to-Strong Generalization cites this paper.

Debate Helps Weak-to-Strong Generalization Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T17:50:56.473929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:50:56.473929Z digest=sha256:4aafeb34e5392fcb3548dff45c2db1aa71ec6818301591d27623a9e1bd5b371d

Observation cc797453-adac-4bdb-ab5e-8508b2d32668 · inbound

Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment cites this paper.

Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T19:56:28.755116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:56:28.755116Z digest=sha256:26e51d952d7095ce4a32f6a82d982926da439b57e62da88d0d37fd2b8d63b234

Observation 654159bd-80c1-483a-a5a9-817426e5da63 · inbound

Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search cites this paper.

Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:07:30.581082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T04:06:23.521344Z digest=sha256:2e0d810bcd1a7aec6e646d5b70a0b4b80714389ddb6eb3c25ff7823520411f7c

Observation f5dbe9b9-0893-4154-bab4-d4f8e1d9ca7d · inbound

The Differences Between Direct Alignment Algorithms are a Blur cites this paper.

The Differences Between Direct Alignment Algorithms are a Blur Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:52:29.541271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-23T03:50:03.720389Z digest=sha256:ecbe34300b53bfbe70f7312ba08eca04bc26120bab2ffc8565baca96cc4e83d3

Observation b8af2d07-08a3-4725-832d-7f275034c4bd · inbound

On Teacher Hacking in Language Model Distillation cites this paper.

On Teacher Hacking in Language Model Distillation Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T11:34:57.128745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:34:57.128745Z digest=sha256:2e19b24290fc6830f250406baf96387695bbf414bf95ac0aed09b4308a2de2e1

Observation 5f84b809-9c0b-464b-9ae6-7e16dff56c16 · inbound

Design Considerations in Offline Preference-based RL cites this paper.

Design Considerations in Offline Preference-based RL Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.034684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.034684Z digest=sha256:3efeb95b4e92c14d65ada5112d850a442e11ebecbb249301dd64a01c7017cda7

Observation 85b9a4a5-97aa-499d-8585-18d3b2556cd4 · inbound

Preference learning made easy: Everything should be understood through win rate cites this paper.

Preference learning made easy: Everything should be understood through win rate Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.165463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.165463Z digest=sha256:a81da8776aef6787ec36e060bd5ed5977d6ff167a414d8f280dd204b3a917d07

Observation 2e5be1d0-c8d3-4d4d-b9bd-fb63ec9bd4c0 · inbound

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation cites this paper.

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:24:12.951783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T07:24:12.845841Z digest=sha256:01adfb79a1d7d1126f375796b3bfbc785f1cf89e85ef2186f9d850f8e5b593a0

Observation 12ba28b3-adb4-40ea-813f-c0a3460688a0 · inbound

Generative AI Act II: Test Time Scaling Drives Cognition Engineering cites this paper.

Generative AI Act II: Test Time Scaling Drives Cognition Engineering Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 282

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:43.718310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:02:43.718310Z digest=sha256:991f854c1eee8c2b72478b2779df74c853ab2652f1768b811453d3a1b972f0e5

Observation c50e7a09-cd8c-4f95-834e-d184db24cd46 · inbound

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function cites this paper.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:11.242520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:11.242520Z digest=sha256:80c56f94f2cbfe6ca54abfeb4e52a1908d08888346234e620cea6d37e009fa6f

Observation 14082912-204e-42f1-be50-db845eab2700 · inbound

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints cites this paper.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.735489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.735489Z digest=sha256:762d5f724484ea8357b0d3d57a3865571ab61d9fd449ce55a90f84b547803aab

Observation e92e82e7-7810-42a0-a808-f87d2fe05105 · inbound

MotionFlux: Efficient Text-Guided Motion Generation through Rectified Flow Matching and Preference Alignment cites this paper.

MotionFlux: Efficient Text-Guided Motion Generation through Rectified Flow Matching and Preference Alignment Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 497

Resolution
unresolved
no resolver link, observed 2026-08-15T16:56:25.196345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:56:25.196345Z digest=sha256:61a636be0bc5d9f9ae2d58e6a923321cae82b4f96830c2c813e7f9c5b98c7305

Observation 3792cc63-5474-4262-842a-4b7510ffd3c7 · inbound

Failure Modes of Maximum Entropy RLHF cites this paper.

Failure Modes of Maximum Entropy RLHF Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:02:39.742787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-18T14:02:11.084514Z digest=sha256:4906dca9cb2d70d85a77cd7f63e88e7a0d1e28c3f6a5f836afaf9650de6e2de9

Observation 3b0f04f9-5a26-44f5-aa62-e5839e6ae26b · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:21.625861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:21.625861Z digest=sha256:6f4efa2d5a98f71ce4aba22f295f6a47ef505061c7d17134155b0cf845c15ff3

Observation c9a5dd00-2d51-4200-862c-00d96d761481 · inbound

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses cites this paper.

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:58.777308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T01:13:29.292351Z digest=sha256:14251851af79981988c08a59933e77444465305134482835d4d16631fde80593

Observation c85fdfea-fd70-4bc8-876d-ca122f1d053c · inbound

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities cites this paper.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:22:41.901099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:56942b1626daad0a38883dfc180c3fd20c19d19f770948d403e7410476fe4346

Observation f77a5d46-b854-482b-a836-d89b4262c404 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 142

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.190388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:3888ba0d73305979b5c919e605c7ae26a0f0b2dd3253d5614b2e818b243bdc14

Observation 052f7ede-3118-4acd-924c-a2163db2c137 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 142

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:45:06.584220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:82863d33892144d0c6745061ed56cb0115309f4df2653aef06addd9459abd437

Observation c8f2e35b-8718-4d47-a563-965679ca8cff · inbound

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges cites this paper.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-08T21:25:38.651895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:6ca53b0d7718e7a6cf2ba82a77373608fea79aae17db4d78780b8f7d42eb72ff

Observation c101aeed-a13c-4375-bc23-474edfb88b27 · inbound

Rater State Bias in RLHF Preference Data: An Audit Framework cites this paper.

Rater State Bias in RLHF Preference Data: An Audit Framework Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T16:22:43.863555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:22:43.863555Z digest=sha256:392c0bafa46afb4a9c6634440b9cdc4c4e10b16ad91bcb062dd29594ac2c4bcf