Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2406.02900.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:43.718310Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T21:25:38.650008Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 3ed7ad6e-b5db-4026-8e0b-6a6ffabe27d5 · inbound
Efficient Alignment of Large Language Models via Data Sampling Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6113cc10-fc17-433c-ba21-e4283a48f874 · inbound
Does RLHF Scale? Exploring the Impacts From Data, Model, and Method Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c46f08b-c16d-4d5b-b3dc-097a730d8a2d · inbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9ef3da2-eead-474a-8b86-200a858c4dc4 · inbound
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea404f9e-bb08-4250-989e-7ef428c33648 · inbound
AlphaPO: Reward Shape Matters for LLM Alignment Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb2d9cc5-83f4-4ffb-a3af-d78b90e26c4c · inbound
T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc65dccf-c32c-4e79-840d-5f81612232a5 · inbound
Online Preference Alignment for Language Models via Count-based Exploration Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa73abb2-54c2-4e48-a89f-a33772d9189b · inbound
Debate Helps Weak-to-Strong Generalization Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc797453-adac-4bdb-ab5e-8508b2d32668 · inbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 654159bd-80c1-483a-a5a9-817426e5da63 · inbound
Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f5dbe9b9-0893-4154-bab4-d4f8e1d9ca7d · inbound
The Differences Between Direct Alignment Algorithms are a Blur Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b8af2d07-08a3-4725-832d-7f275034c4bd · inbound
On Teacher Hacking in Language Model Distillation Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f84b809-9c0b-464b-9ae6-7e16dff56c16 · inbound
Design Considerations in Offline Preference-based RL Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85b9a4a5-97aa-499d-8585-18d3b2556cd4 · inbound
Preference learning made easy: Everything should be understood through win rate Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e5be1d0-c8d3-4d4d-b9bd-fb63ec9bd4c0 · inbound
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 12ba28b3-adb4-40ea-813f-c0a3460688a0 · inbound
Generative AI Act II: Test Time Scaling Drives Cognition Engineering Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 282
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c50e7a09-cd8c-4f95-834e-d184db24cd46 · inbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14082912-204e-42f1-be50-db845eab2700 · inbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e92e82e7-7810-42a0-a808-f87d2fe05105 · inbound
MotionFlux: Efficient Text-Guided Motion Generation through Rectified Flow Matching and Preference Alignment Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 497
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3792cc63-5474-4262-842a-4b7510ffd3c7 · inbound
Failure Modes of Maximum Entropy RLHF Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3b0f04f9-5a26-44f5-aa62-e5839e6ae26b · inbound
Adaptive Margin RLHF via Preference over Preferences Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9a5dd00-2d51-4200-862c-00d96d761481 · inbound
$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c85fdfea-fd70-4bc8-876d-ca122f1d053c · inbound
Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f77a5d46-b854-482b-a836-d89b4262c404 · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 142
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 052f7ede-3118-4acd-924c-a2163db2c137 · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 142
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c8f2e35b-8718-4d47-a563-965679ca8cff · inbound
More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c101aeed-a13c-4375-bc23-474edfb88b27 · inbound
Rater State Bias in RLHF Preference Data: An Audit Framework Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.