Pith. sign in

Paper Citation Record · LEDGER

Direct Advantage Regression: Aligning LLMs with Online AI Reward

As of 19 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2504.14177.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.14177 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:00:34.711691Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f67c0f3a-d532-48a4-a3dc-e0644bbc088d · outbound

This paper cites Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms, 2024.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:00:35.951192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T12:00:34.204464Z digest=sha256:b802db67448eb9cd35e6fc9a904d102c995a53974873527bd53d0f2c003e83a7

Observation 78f90cf9-0683-4843-9857-75deac2d3d4d · outbound

This paper cites A General Theoretical Paradigm to Understand Learning from Human Preferences.

Direct Advantage Regression: Aligning LLMs with Online AI Reward A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.211481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.211481Z digest=sha256:f457c999be5f2094332a678db0289ac9a553079ee306cf136d13b4920d00e542

Observation 146d267f-f2b3-4cd1-b680-36f4a59acfc3 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.217863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.217863Z digest=sha256:7b2f62c3db9acf95e53d0bf6cb03f6a7e2ec2b5149812ae3e99de8386905c2cf

Observation f0cc1340-df53-43f0-b848-1762fafa1f8f · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.223434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.223434Z digest=sha256:8a28f724f75fa8eebd51cac6193dc0f55945172096edca26cae8c3348bd187f2

Observation b3641164-07a0-4499-ac7b-ed9b8202db4a · outbound

This paper cites Language Models are Few-Shot Learners.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.229919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.229919Z digest=sha256:e17ee29558f4fecc10c804f9e9777c64b6bbff1f934c9db46a44227024e5059c

Observation 64188103-c8e8-4033-ac65-247f4f919641 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.236127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.236127Z digest=sha256:2c68c5b14c38a347c8f4509cfe6a58d68ccc61b83f2d05e5a9eb909955d9306a

Observation 3c207f77-5e88-462b-9a07-a23162186ff2 · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.245263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.245263Z digest=sha256:f0fdf1a787110a020d5d83d4ad9e7b279882b422edd362f466fea48d683b5c22

Observation 62a67080-c3a3-4199-9a4a-e2c24cd6fc9a · outbound

This paper cites Bootstrapping Language Models with DPO Implicit Rewards.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Bootstrapping Language Models with DPO Implicit Rewards

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.252305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.252305Z digest=sha256:6d4c0ada7d7b072bb3221420fd7d3597f5b79a6e8075c4845b27891bffc69b46

Observation 1ee19758-c12f-4241-8e2c-22849ea4afd8 · outbound

This paper cites OPTune: Efficient Online Preference Tuning.

Direct Advantage Regression: Aligning LLMs with Online AI Reward OPTune: Efficient Online Preference Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.259371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.259371Z digest=sha256:c5f550de49fd6c70ceb7fe9d166cf1a4c69581190b5f296c04c36dd8104e1770

Observation b1a6ddfb-930d-43be-9f07-d598e35b906a · outbound

This paper cites GRATH: Gradual Self-Truthifying for Large Language Models.

Direct Advantage Regression: Aligning LLMs with Online AI Reward GRATH: Gradual Self-Truthifying for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.274134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.274134Z digest=sha256:7e107ff9f5cc1c64e0387b12a4ae12a82ffb865171633e62523e37341092995c

Observation 79bee758-fb04-48c2-b7e9-39f29f96472d · outbound

This paper cites Cost-Effective Proxy Reward Model Construction with On-Policy and Active Learning.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Cost-Effective Proxy Reward Model Construction with On-Policy and Active Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.281660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.281660Z digest=sha256:0253cc22be52cca413b422127291aff3573a38e7d289afff3ff743cfa3b9ebd6

Observation 47a838fc-ef96-4b8e-b869-677ec81c36e4 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Direct Advantage Regression: Aligning LLMs with Online AI Reward UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.290438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.290438Z digest=sha256:4d2660cb37ad14ea2b6d8e416859238f51e1fa3e7ae060bb36c9ae3db8c5f613

Observation 88d9a818-1973-4e64-9c2c-5823feadb2e0 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Direct Advantage Regression: Aligning LLMs with Online AI Reward FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.296889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.296889Z digest=sha256:6d0e624ecab33bc3c96d6f1c701a97682dd3dcfb0ab6ae86b2e64f393b2ef407

Observation ab6b6760-305f-4e42-bc1e-e73676c10f70 · outbound

This paper cites The Llama 3 Herd of Models.

Direct Advantage Regression: Aligning LLMs with Online AI Reward The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.303916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.303916Z digest=sha256:2a02872c26d5957ed171b7d259f3999951d7953f83b3adf983e0fc30b507e0bc

Observation b14e1d8f-a538-42e8-bc4e-bb8f264f61b7 · outbound

This paper cites Accelerate: Training and inference at scale made simple, efficient and adaptable.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Accelerate: Training and inference at scale made simple, efficient and adaptable

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.309446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.309446Z digest=sha256:3209db723fc211c30b970f25d9dea0a675f0fa37d78372e1239849ec9cab41c1

Observation ca4695f6-5b45-4b13-ba43-bf7e576d529a · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.315247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.315247Z digest=sha256:5d1d9673496bfd7041a47d00164fe1d2d3226a67b7a42d356a889e5164349498

Observation b53c1a16-f0ec-4732-9240-220720a77ba0 · outbound

This paper cites Large Language Models Can Self-Improve.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Large Language Models Can Self-Improve

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.320360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.320360Z digest=sha256:26681b56ab9299329afa94e62b374a2b821d993aedc59db5b4167df7534ccf1b

Observation 5bb3867a-9e85-4ccd-9a0b-dd040d95acce · outbound

This paper cites Mistral 7B.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Mistral 7B

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.327473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.327473Z digest=sha256:ea8c7850683f5969002bb85ee21b6bc1e5cf8b11d16f2a9b9103ad560432b946

Observation fb1df4bc-5fd6-47c1-9bbb-0fdf51c42bb1 · outbound

This paper cites LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression.

Direct Advantage Regression: Aligning LLMs with Online AI Reward LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.332928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.332928Z digest=sha256:4d6b35bc6c347e4dae3274eabebf9b0e1b9eb86281c53dc941f5a2f995632f64

Observation 49d30676-1d5e-4e45-9221-f88207877ce6 · outbound

This paper cites Buy 4 REINFORCE samples, get a baseline for free!, 2019.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Buy 4 REINFORCE samples, get a baseline for free!, 2019

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:00:35.911831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T12:00:34.340463Z digest=sha256:715488d61ae82c5dab48ddd025f0c3b75cd131993bad40fe800e3db4c5e857c7

Observation 7db09f20-10ba-4ae2-95c2-932de7832e3b · outbound

This paper cites H., Gonzalez, J.

Direct Advantage Regression: Aligning LLMs with Online AI Reward H., Gonzalez, J

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.347155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.347155Z digest=sha256:0daebc2e2f3ccad267befe9caf118dee8dbf480e0e9fe5d0da35f5ee0ac92b9f

Observation 5713cc60-a032-40f9-9ee1-9c7fb09b7853 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Direct Advantage Regression: Aligning LLMs with Online AI Reward RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.355051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.355051Z digest=sha256:929242501033ed21908bda81e790fe25c22d50d0f381c3efffd7af353c68aa97

Observation 119025eb-6581-4f30-876e-c57e1f9bbfa5 · outbound

This paper cites Long-context LLMs Struggle with Long In-context Learning.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Long-context LLMs Struggle with Long In-context Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.363013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.363013Z digest=sha256:dcbf6eddbecb4982d89912adeec1488590cef8f26e8ded0672dc85846bf13d41

Observation 0260710d-1281-41b8-94ff-f9cfb170da8d · outbound

This paper cites Self-Alignment with Instruction Backtranslation.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Self-Alignment with Instruction Backtranslation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.370186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.370186Z digest=sha256:45140c7f099d0f038ce77d9fbbd03e0c0a6933a6fdbc53938eee20c424cbe0de

Observation 655e60ba-9d44-4f89-93ce-da6d3994aeb9 · outbound

This paper cites GPT-4 Technical Report.

Direct Advantage Regression: Aligning LLMs with Online AI Reward GPT-4 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.376535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.376535Z digest=sha256:28d95e8046ffa8f94f6696d19384bbe027342cda97a46ae8382808de494d1a7e

Observation 2823e698-7e03-4f77-ab70-9a1cd46d21df · outbound

This paper cites Training language models to follow instructions with human feedback.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Training language models to follow instructions with human feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.383490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.383490Z digest=sha256:20faf8f97261b94371598384d8716b9314d1507b34aa19dabff40a95f7afb0e1

Observation beaa816a-bc9f-4c5b-9e67-e5969ca58a1c · outbound

This paper cites Iterative Reasoning Preference Optimization.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Iterative Reasoning Preference Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.388848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.388848Z digest=sha256:85c3602cc83d5ec7258dc2fe07f97a0a0b9859f9c362f69d5261b78e72d79b12

Observation 567618af-c880-49f7-8ed0-53554abed509 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.396199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.396199Z digest=sha256:bd69b745ee158d79a29196311be740d98685d055a7e3b2dcab70a013da2297f7

Observation 566be659-2760-4af5-aca0-72c8a17bb822 · outbound

This paper cites and Schaal, S.

Direct Advantage Regression: Aligning LLMs with Online AI Reward and Schaal, S

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.401828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.401828Z digest=sha256:313e8dc3bfad54a8e5435451e60e93f382df14bb45b348304b3f5787294d0602

Observation 7d34b619-f8f7-451f-83af-f53ae90e55a7 · outbound

This paper cites Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.407637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.407637Z digest=sha256:67b6174776889ab9e351b91a3db3314d0487e1da84cff6a91d87630e3a4dbecd

Observation 2dbdcd2e-4fe1-4134-8f3a-a9a247fe76f6 · outbound

This paper cites Improving language understanding by generative pre-training.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Improving language understanding by generative pre-training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.417554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.417554Z digest=sha256:cf4b2e6f7d380dce63b9a665254b5c079948895fb3200b61cf5c54b2630d1bf2

Observation 2451e565-9a03-44d8-a62f-5a9877d51b31 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.423837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.423837Z digest=sha256:91fdb84a210db9eacf1660c08c7ac93718762c93adda3ce8730d56fc2a8eeaa4

Observation bd5f79f4-7d37-4d7b-a39b-23b9227ab623 · outbound

This paper cites Stable-baselines3: Reliable reinforcement learning implementations.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Stable-baselines3: Reliable reinforcement learning implementations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.430223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.430223Z digest=sha256:09cdf698adc2d17bd570ae63a958fa490a7d86a310d16a6131948f465d1dd9f3

Observation 80c2df79-a520-471c-99a7-e48dcffd2c15 · outbound

This paper cites Trust Region Policy Optimization.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Trust Region Policy Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.437502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.437502Z digest=sha256:1bb974b5e04daeb13c24ec91cb0cd6eea1d618166c70aba0f097f38aa2feb3b7

Observation 187f42c2-3341-4438-8a03-2b8e6a2d0b79 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.444215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.444215Z digest=sha256:7efc67a56de790a99f3875176569b694f408a194c6ee3fcb4150691b2741489f

Observation d9c77d75-1e90-4a04-9c62-155cf302ce07 · outbound

This paper cites Adafactor: Adaptive Learning Rates with Sublinear Memory Cost.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Adafactor: Adaptive Learning Rates with Sublinear Memory Cost

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.449369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.449369Z digest=sha256:b448851244fd657e34a7a38e596cd357c4b78c8630c073d43dde31d45991c79d

Observation e3f413ea-612b-4baa-b9c8-235f2c723be3 · outbound

This paper cites Learning to summarize from human feedback.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Learning to summarize from human feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.459183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.459183Z digest=sha256:12bdd59f44c6d8c8ade14b068fb050935021ba641284c9e33087d75a2bdde404

Observation 54d68c2b-3113-4974-83dc-45c6416b9e11 · outbound

This paper cites S., Barto, A.

Direct Advantage Regression: Aligning LLMs with Online AI Reward S., Barto, A

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:00:35.834940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T12:00:34.466885Z digest=sha256:4b5667643f6b7893fb55d76b6fbc376a8f5e869688d38821b9c356cddb8b0615

Observation 3bfda666-2dd3-41a5-b7b1-ab9c1a87c596 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Gemma 2: Improving Open Language Models at a Practical Size

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.471981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.471981Z digest=sha256:19084276d66499fd2e3ccb565f811e7154994f88c521902a1bdb268c7c580ccc

Observation cff0372b-2ae2-445b-91d2-58f041f27ecf · outbound

This paper cites Attention Is All You Need.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Attention Is All You Need

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.479259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.479259Z digest=sha256:d1399433e9c6d70023a0c0e4c67774a518a95a80c2ba845cd328479645640348

Observation c903f103-4cda-4f66-967b-b0d7a5ae870d · outbound

This paper cites Trl: Transformer reinforcement learning.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Trl: Transformer reinforcement learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.487071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.487071Z digest=sha256:7a13b6e03779b18c8af822ec72a9317497fbfc1390bdd8c29d019494bbe88c7e

Observation b9a77a3b-2132-4a58-8abc-c065d1b2e5c2 · outbound

This paper cites Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.644366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.644366Z digest=sha256:fe657eb36b4af3a60afa3e6984b4273bb1a2cdb0fb802d3c1cd7826d64e2255c

Observation cca40d48-4670-4b1e-bf65-19fca80100bb · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.649935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.649935Z digest=sha256:48d93aa29670a9fb31f1a47ddf6b7b8615784b6664f3d2bfbb743a21fd291c01

Observation ea32d943-8abb-4c55-b81a-45742f6594ab · outbound

This paper cites Critic Regularized Regression.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Critic Regularized Regression

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.655432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.655432Z digest=sha256:03eac3a60798975dafa3daf913432aba55f3db5fd8b87d089f9d69d9da6e4577

Observation 851b486b-a22a-46a7-a95b-1d03bfab8819 · outbound

This paper cites HelpSteer2-Preference: Complementing Ratings with Preferences.

Direct Advantage Regression: Aligning LLMs with Online AI Reward HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.661542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.661542Z digest=sha256:a1f361bd65ef68e4255332882934c70cb6c09a28bc009bd9575deffc9de96cb3

Observation 08cb2999-9b56-480d-a747-9b1a019abc15 · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.666679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.666679Z digest=sha256:934da77ea721f9a8b1b77a541ef63ba63c1c077105701cca49ea2b26bff2ab7e

Observation e83acff7-a118-4e3e-a419-bded8c159674 · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.671758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.671758Z digest=sha256:c929a40ee333d5ab3bdc7a57e68ff53c013b2e15fa565fd9512deca6c128f004

Observation 0ac64dd9-830b-43d8-9486-62a9a033e002 · outbound

This paper cites Qwen2 technical report, 2024 a.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Qwen2 technical report, 2024 a

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:00:35.799127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T12:00:34.677912Z digest=sha256:2ca79efdfa15a0ddd761c129247f9cdcd201eb0854b650ee7871ca8c5bf7a157

Observation f89034d2-5489-42bd-b99b-6761bad044ed · outbound

This paper cites Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.683581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.683581Z digest=sha256:aa740903970d24828631e1630b791d863571d48d4fd05f817b5fa8e86e6871d4

Observation 86d7ba9b-d3bb-4c55-9b34-b70436a2615a · outbound

This paper cites Self-Rewarding Language Models.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Self-Rewarding Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.690216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.690216Z digest=sha256:a72b490c22c9af50e3cd157e232553516f73e2bcc32f9d88940b005f34e763ea

Observation e9883153-07d2-4036-bb4c-bde0e9152d9f · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Direct Advantage Regression: Aligning LLMs with Online AI Reward SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.695729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.695729Z digest=sha256:68c418e431bf0639c149d10b6cfeec112cff5b2c3e3cecdf4ea9ab6859bb3e6e

Observation a9e70935-7f50-4f1e-8e40-9fc46f99f64c · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.701059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.701059Z digest=sha256:facfdabaa253b0cb64a0a87a86652458df760491bd8e1dee4c0f855f87903ebf

Observation b71ae924-5248-4a6e-8a29-e29532d92614 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Direct Advantage Regression: Aligning LLMs with Online AI Reward Fine-Tuning Language Models from Human Preferences

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.706600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.706600Z digest=sha256:f21e1eda5183d0192d218823c115cfacc94c9d72d15dec5b4d085f89502aeb9f

Observation 8711beae-a82d-419a-a3f5-a8e0c68b3288 · outbound

This paper cites write newline.

Direct Advantage Regression: Aligning LLMs with Online AI Reward write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:34.711691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:00:34.711691Z digest=sha256:bf51d554524b9a7c7a70d9975b2070c407625d08f7e2a5dda80f5e50e7bc615d

Pith citing papers

No inbound Pith citation observations are available.