Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:40:18.045861Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2506.07492.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:40:18.045861Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T16:11:49.149334Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T09:11:01.276341Z
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 814c293d-8351-4c84-a266-071dd259de96 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7733375b-fae1-45c0-94d4-4195aa3d2497 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28a73e90-003e-4e7b-b6da-7f358d3c3f51 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Llama 3 model card
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 587bad29-1ff8-4a50-9e75-684abf84c47f · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Direct Preference Optimization with an Offset
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79c6ad72-af3f-4126-a1aa-0601e6c05547 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model G., Guo, Z
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c395d6de-c621-4c82-81ed-498728b3de65 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 946f90a8-9e91-4dea-b308-9bf73f94ebfa · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Constitutional AI: Harmlessness from AI Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99bb216e-2624-4f7b-be33-c84082939d0b · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0ef421c-5e0e-41ef-a713-4ec9ce79430c · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model and Rinaldo, A
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a615d5e9-8a6f-43da-84cd-91beb628846a · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0320c2d-dee5-4d5f-b887-0426e3b56aea · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4011480-7ab6-4093-8a94-ac2d90a9758d · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model A survey on evaluation of large language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 48ce16f0-6567-4693-b4f1-353b408a368f · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Bootstrapping Language Models with DPO Implicit Rewards
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bcc5977-6d0d-40de-9a89-38f77c56e04e · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Ultrafeedback: Boosting language models with scaled ai feedback
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23d7df53-4660-4f6d-8841-aae19e1168d8 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Enhancing chat language models by scaling high-quality instructional conversations
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3d3c575-fa88-4367-a5b8-4c98693dcc9f · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7b3bf60-134f-42c1-a96a-de45a050d49f · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model KTO: Model Alignment as Prospect Theoretic Optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29c4a37d-1f34-468f-90e3-68cb338f922a · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ac79af2-8955-4e80-8e72-19b881aba6d1 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Bias and Fairness in Large Language Models: A Survey
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee9c6377-7ddb-4e9b-8ca8-df6af4d58ad9 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bf454c0-d0c9-470b-99ac-b9b245bc4f18 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Learn Your Reference Model for Real Good Alignment
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 186c1af7-8365-4919-945b-3665b82d0bcd · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model and Pierskalla, W
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 69ecbe00-bcf7-4428-91ab-5fbd69ed5fa4 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e897ace9-7697-498a-ba60-dd00acfaa9c8 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model ORPO: Monolithic Preference Optimization without Reference Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1a6ffa6-1c87-438c-9b8e-5c3b6a11a067 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Understanding the Learning Dynamics of Alignment with Human Feedback
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6a10d37-3832-4f5e-b2e4-bf3e7bf18d1e · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model https://github.com/huggingface/trl/pull/1265
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65a14ff0-d858-4efd-8bd0-57f9600c6629 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Adam: A Method for Stochastic Optimization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dba99e6-d37a-40a8-8f2b-710645e6e6eb · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Common learning constraints alter interpretations of direct preference optimization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06e98cfe-841e-411d-b3be-83c479524cec · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model H., Gonzalez, J
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9daa881e-8c3b-4878-b8ce-b212ac68cc8f · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7c6d022-71b5-42c7-90d9-c7b5bcbd23f7 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Policy Optimization in RLHF: The Impact of Out-of-preference Data
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08f0ace7-0ead-474d-aaaf-3a7fb6b1e2fe · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9735fcff-1b7a-4138-b7d9-0ee2d8104416 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model L., Daly, R
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3eaad74c-b477-4f61-ad2d-2514ce7af7ab · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4d29e7f-976f-40d7-9278-28a683cbe2cb · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Active Preference Learning for Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31a19b43-e3f6-46ee-8746-81b6e9241408 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Training language models to follow instructions with human feedback
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59fc338f-5672-41ee-81f7-61e83c74674c · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2444c89a-c34f-44c2-a429-af7a17a9eb79 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Disentangling Length from Quality in Direct Preference Optimization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 491e4e82-0774-4f99-87e1-63e182a5feaa · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aa11bc4-da3e-4fd0-9bad-4b0f8063e226 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model and Schaal, S
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b783b59-d2c2-4763-b8b5-40991e1dd7b1 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model D., Ermon, S., and Finn, C
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4315a4d-6138-4a5e-a77a-56c8f4f8748a · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a20d07-a754-48c3-9370-5e3d676483f5 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62dbb63e-f8b3-434a-8ba4-510035797112 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Proximal Policy Optimization Algorithms
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f0c31c3-9de9-486f-a220-7e3369074278 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e8ed1e2-e9a0-46cb-861c-81de59126bae · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model The importance of online data: U nderstanding preference fine-tuning via coverage
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf476f93-313a-4550-9249-1b0ca218ec18 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab772d2f-269f-40f3-bd2f-4188f6b385f7 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model S., and Bagnell, J
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c06bc6f8-2862-4405-a46a-971305cb837d · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bf9cade-a35f-49df-802a-f4f1d1c79197 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Generalized Preference Optimization: A Unified Approach to Offline Alignment
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b95adbf-7431-45c7-ada4-c04e01e142f6 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Beyond reverse KL : Generalizing direct preference optimization with diverse divergence constraints
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88349501-5533-413c-8a10-33c1c19c004b · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model mDPO: Conditional Preference Optimization for Multimodal Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7cab4d7-a511-4c0c-860c-3f7a5215f708 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3613ae8d-bd79-4c5b-9de9-0ce589cb6c87 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model A Survey of Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96a4e02e-511b-44ed-9d04-e4057629f929 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05aafb46-278e-4a2c-b53c-2bf326be5b86 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model Fine-Tuning Language Models from Human Preferences
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc3da6ce-ceed-4082-a2d3-b4792f7fe476 · outbound
Explicit Preference Optimization: No Need for an Implicit Reward Model write newline
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bb23c37-e426-491f-85e8-4e5cf80e523b · inbound
Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization Explicit Preference Optimization: No Need for an Implicit Reward Model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.