Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:31:49.368412Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.13702.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:31:49.368412Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T14:24:20.906356Z
A source-named dated measurement, never combined with another source.
Source: cited_works
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7d1e438d-cf68-4314-9294-e3aba8e89a37 · outbound
Value-Free Policy Optimization via Reward Partitioning Rrhf: Rank responses to align language models with human feedback,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2be3843c-13a7-4a96-be30-6551e33c15e0 · outbound
Value-Free Policy Optimization via Reward Partitioning RLHF Workflow: From Reward Modeling to Online RLHF
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6740c56-52c9-4774-bd8a-e41a14d51045 · outbound
Value-Free Policy Optimization via Reward Partitioning Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0364f190-1d28-4505-b39d-323425002d42 · outbound
Value-Free Policy Optimization via Reward Partitioning Direct preference optimization: Your language model is secretly a reward model,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 170bbaa3-093d-4e6d-a9ea-1ab59660b784 · outbound
Value-Free Policy Optimization via Reward Partitioning Generalized preference optimization: A unified approach to offline alignment,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 848849fa-a948-449d-976a-015b8f7fb801 · outbound
Value-Free Policy Optimization via Reward Partitioning Offline Regularised Reinforcement Learning for Large Language Models Alignment
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffe4e4d3-d6bc-4aa1-8966-aed06214946b · outbound
Value-Free Policy Optimization via Reward Partitioning Model alignment as prospect theoretic optimization,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6699bbdc-bc91-4330-9be9-e3013b84bd57 · outbound
Value-Free Policy Optimization via Reward Partitioning Instruction tuning for large language models: A survey,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb0b05c5-085b-4d2c-b061-a4849a9ca341 · outbound
Value-Free Policy Optimization via Reward Partitioning Proximal Policy Optimization Algorithms
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff63cb40-f6da-4a37-bc7b-9a7815d04dff · outbound
Value-Free Policy Optimization via Reward Partitioning Trust region policy optimization,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b6d7598-5178-4239-ba4a-ae55ba16e2ab · outbound
Value-Free Policy Optimization via Reward Partitioning Training language models to follow instructions with human feedback,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc503428-b3a4-4786-8dab-a4ce942ed59c · outbound
Value-Free Policy Optimization via Reward Partitioning GPT-4 Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a6f198-8d6c-4b8d-bc1c-481155b2fd4c · outbound
Value-Free Policy Optimization via Reward Partitioning The claude 3 model family: Opus, sonnet, haiku,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3056ed27-b9f2-4632-b3e7-ae35c551fc55 · outbound
Value-Free Policy Optimization via Reward Partitioning Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cff56b3c-ca0c-458e-92b1-3cd249e28eca · outbound
Value-Free Policy Optimization via Reward Partitioning Guiding pretraining in reinforcement learning with large language models,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d3ab8c7-82c6-472a-85b1-493e848b096b · outbound
Value-Free Policy Optimization via Reward Partitioning RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e48b80a-f4a5-4b8c-a791-f04e80f8181f · outbound
Value-Free Policy Optimization via Reward Partitioning CREAM: consistency regularized self-rewarding language models,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0267a425-915e-45a6-a8be-60f29fa65428 · outbound
Value-Free Policy Optimization via Reward Partitioning Self-rewarding language models,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 586052fa-c7b1-42f5-b7d7-8e467d6e9ebc · outbound
Value-Free Policy Optimization via Reward Partitioning SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15bae0be-35c2-44a2-bd98-8b0a8a551dbd · outbound
Value-Free Policy Optimization via Reward Partitioning The Llama 3 Herd of Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 163ec618-6992-4fcf-ae7d-625c5dd4cf0a · outbound
Value-Free Policy Optimization via Reward Partitioning Qwen3 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2acf63e6-e1f2-4458-9c6f-427069ec6d62 · outbound
Value-Free Policy Optimization via Reward Partitioning Nemotron-4 340B Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d4fdb80-5a78-4520-b10c-85b41aeb9bdf · outbound
Value-Free Policy Optimization via Reward Partitioning Rank analysis of incomplete block designs: I. the method of paired comparisons,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0633016-44d7-4956-a4d3-a56c310edf3e · outbound
Value-Free Policy Optimization via Reward Partitioning A general theoretical paradigm to understand learning from human preferences,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f496c456-0247-4c12-b42d-338ec07bae57 · outbound
Value-Free Policy Optimization via Reward Partitioning Advances in prospect theory: Cumulative representation of uncertainty,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation efb414d2-4669-4b7d-a5eb-e8b65446ba0b · outbound
Value-Free Policy Optimization via Reward Partitioning UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7262b461-5e87-474b-85dd-4960bc1d8d80 · outbound
Value-Free Policy Optimization via Reward Partitioning Alpacaeval: An automatic evaluator of instruction-following models,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 29968375-32d7-49df-bf73-acb155c2e158 · outbound
Value-Free Policy Optimization via Reward Partitioning Judging llm-as-a-judge with mt-bench and chatbot arena,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b78be03-5ee5-44bc-9e33-ffbaca2aee97 · outbound
Value-Free Policy Optimization via Reward Partitioning Instruction-Following Evaluation for Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b07c684-5b46-4e78-b412-79e35815b2c1 · outbound
Value-Free Policy Optimization via Reward Partitioning Training Verifiers to Solve Math Word Problems
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17fce266-5c00-4cb7-852f-8059491a6f1c · outbound
Value-Free Policy Optimization via Reward Partitioning Scaling up models and data with t5x and seqio,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7fb1bfa7-74ea-4b4e-93ff-1c7f6ca8f999 · outbound
Value-Free Policy Optimization via Reward Partitioning Mistral 7b,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab1ceb4b-8ef4-47cf-9ba4-9e85d1442e9d · outbound
Value-Free Policy Optimization via Reward Partitioning Llama: Open and efficient foundation language models,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation edba6570-de0e-4a9b-9180-4ac294105a90 · outbound
Value-Free Policy Optimization via Reward Partitioning Qwen2 Technical Report
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 526d418b-dd00-4c0e-8f9d-32f81c3ca79a · outbound
Value-Free Policy Optimization via Reward Partitioning BERTScore: Evaluating Text Generation with BERT
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 762e49c1-11ac-4aac-bf6e-ea267e0d46be · outbound
Value-Free Policy Optimization via Reward Partitioning Rouge: A package for automatic evaluation of summaries,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c2fe268-4e4a-4beb-85cc-8fae2cd7f522 · outbound
Value-Free Policy Optimization via Reward Partitioning A diversity-promoting objective function for neural conversation models,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c84dfb56-bb9d-4b13-8610-5d333745a163 · outbound
Value-Free Policy Optimization via Reward Partitioning A Survey on LLM-as-a-Judge
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c8e74f6-81cf-4fd3-9c7a-df18274cf243 · outbound
Value-Free Policy Optimization via Reward Partitioning GPT-4o System Card
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8c59ea9-5c99-45a5-bbd7-2c6f3921eecc · outbound
Value-Free Policy Optimization via Reward Partitioning Claude 3.5 sonnet
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a806028-570f-490e-bdc4-e44b439ba7ac · outbound
Value-Free Policy Optimization via Reward Partitioning Decoupled weight decay regularization,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58037a65-8d50-481a-8f54-57238aecf6dd · inbound
Safety Alignment of LMs via Non-cooperative Games Value-Free Policy Optimization via Reward Partitioning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.