Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:00:34.711691Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2504.14177.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:00:34.711691Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f67c0f3a-d532-48a4-a3dc-e0644bbc088d · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 78f90cf9-0683-4843-9857-75deac2d3d4d · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 146d267f-f2b3-4cd1-b680-36f4a59acfc3 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0cc1340-df53-43f0-b848-1762fafa1f8f · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Constitutional AI: Harmlessness from AI Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3641164-07a0-4499-ac7b-ed9b8202db4a · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Language Models are Few-Shot Learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64188103-c8e8-4033-ac65-247f4f919641 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c207f77-5e88-462b-9a07-a23162186ff2 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62a67080-c3a3-4199-9a4a-e2c24cd6fc9a · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Bootstrapping Language Models with DPO Implicit Rewards
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ee19758-c12f-4241-8e2c-22849ea4afd8 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward OPTune: Efficient Online Preference Tuning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1a6ddfb-930d-43be-9f07-d598e35b906a · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward GRATH: Gradual Self-Truthifying for Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79bee758-fb04-48c2-b7e9-39f29f96472d · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Cost-Effective Proxy Reward Model Construction with On-Policy and Active Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a838fc-ef96-4b8e-b869-677ec81c36e4 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88d9a818-1973-4e64-9c2c-5823feadb2e0 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab6b6760-305f-4e42-bc1e-e73676c10f70 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b14e1d8f-a538-42e8-bc4e-bb8f264f61b7 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Accelerate: Training and inference at scale made simple, efficient and adaptable
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca4695f6-5b45-4b13-ba43-bf7e576d529a · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Direct Language Model Alignment from Online AI Feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b53c1a16-f0ec-4732-9240-220720a77ba0 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Large Language Models Can Self-Improve
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bb3867a-9e85-4ccd-9a0b-dd040d95acce · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Mistral 7B
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb1df4bc-5fd6-47c1-9bbb-0fdf51c42bb1 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49d30676-1d5e-4e45-9221-f88207877ce6 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Buy 4 REINFORCE samples, get a baseline for free!, 2019
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7db09f20-10ba-4ae2-95c2-932de7832e3b · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward H., Gonzalez, J
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5713cc60-a032-40f9-9ee1-9c7fb09b7853 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 119025eb-6581-4f30-876e-c57e1f9bbfa5 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Long-context LLMs Struggle with Long In-context Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0260710d-1281-41b8-94ff-f9cfb170da8d · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Self-Alignment with Instruction Backtranslation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 655e60ba-9d44-4f89-93ce-da6d3994aeb9 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward GPT-4 Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2823e698-7e03-4f77-ab70-9a1cd46d21df · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Training language models to follow instructions with human feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beaa816a-bc9f-4c5b-9e67-e5969ca58a1c · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Iterative Reasoning Preference Optimization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 567618af-c880-49f7-8ed0-53554abed509 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 566be659-2760-4af5-aca0-72c8a17bb822 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward and Schaal, S
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d34b619-f8f7-451f-83af-f53ae90e55a7 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dbdcd2e-4fe1-4134-8f3a-a9a247fe76f6 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Improving language understanding by generative pre-training
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2451e565-9a03-44d8-a62f-5a9877d51b31 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd5f79f4-7d37-4d7b-a39b-23b9227ab623 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Stable-baselines3: Reliable reinforcement learning implementations
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80c2df79-a520-471c-99a7-e48dcffd2c15 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Trust Region Policy Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 187f42c2-3341-4438-8a03-2b8e6a2d0b79 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Proximal Policy Optimization Algorithms
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9c77d75-1e90-4a04-9c62-155cf302ce07 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3f413ea-612b-4baa-b9c8-235f2c723be3 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Learning to summarize from human feedback
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54d68c2b-3113-4974-83dc-45c6416b9e11 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward S., Barto, A
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3bfda666-2dd3-41a5-b7b1-ab9c1a87c596 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Gemma 2: Improving Open Language Models at a Practical Size
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cff0372b-2ae2-445b-91d2-58f041f27ecf · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Attention Is All You Need
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c903f103-4cda-4f66-967b-b0d7a5ae870d · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Trl: Transformer reinforcement learning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9a77a3b-2132-4a58-8abc-c065d1b2e5c2 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca40d48-4670-4b1e-bf65-19fca80100bb · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Self-Instruct: Aligning Language Models with Self-Generated Instructions
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea32d943-8abb-4c55-b81a-45742f6594ab · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Critic Regularized Regression
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 851b486b-a22a-46a7-a95b-1d03bfab8819 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward HelpSteer2-Preference: Complementing Ratings with Preferences
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08cb2999-9b56-480d-a747-9b1a019abc15 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e83acff7-a118-4e3e-a419-bded8c159674 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ac64dd9-830b-43d8-9486-62a9a033e002 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Qwen2 technical report, 2024 a
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f89034d2-5489-42bd-b99b-6761bad044ed · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86d7ba9b-d3bb-4c55-9b34-b70436a2615a · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Self-Rewarding Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9883153-07d2-4036-bb4c-bde0e9152d9f · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9e70935-7f50-4f1e-8e40-9fc46f99f64c · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b71ae924-5248-4a6e-8a29-e29532d92614 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward Fine-Tuning Language Models from Human Preferences
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8711beae-a82d-419a-a3f5-a8e0c68b3288 · outbound
Direct Advantage Regression: Aligning LLMs with Online AI Reward write newline
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.