Pith. sign in

Paper Citation Record · LEDGER

DPO Meets PPO: Reinforced Token Optimization for RLHF

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 39 inbound Pith citation observations for arXiv:2404.18922.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.18922 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 39 of 39 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:13:53.739536Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:26:29.891495Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 66dc5680-099b-439a-8e2a-4c969731c125 · inbound

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution cites this paper.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.968647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.968647Z digest=sha256:e698b8e8897d2beec9ee1586af1f3c721a76fbdfebfb29c1d8ce5dee7a32f70b

Observation 948fde31-2f2d-4ef5-8106-19293324a369 · inbound

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment cites this paper.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.066939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.066939Z digest=sha256:6f2f72c023a023155171bd80b23d90a42e3641793506cedf38a6f31783e0df2f

Observation 96bb879e-6256-4140-988c-471195b71c1d · inbound

T-REG: Preference Optimization with Token-Level Reward Regularization cites this paper.

T-REG: Preference Optimization with Token-Level Reward Regularization DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:56.476509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:15:56.476509Z digest=sha256:6bedd745d5c27f33414ab2c0668ffb861b3c5cb62e65e9c09988c0b291aea74d

Observation f4659949-2b58-4b5b-8532-ade787f3a012 · inbound

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment cites this paper.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.895081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.895081Z digest=sha256:9c5f52dbcef7828f8b72ae275ed9f316ad151c965a3386817d0312e7cf627784

Observation cac1b905-4782-4861-851b-c2214bd75009 · inbound

Online Learning from Strategic Human Feedback in LLM Fine-Tuning cites this paper.

Online Learning from Strategic Human Feedback in LLM Fine-Tuning DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T10:21:33.365940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:21:33.365940Z digest=sha256:8de1b362277509f194f6a6bf65d4232b03a5d5891233aca1195242cc94100cfb

Observation 6b2af11e-e5ae-4fd5-a7ee-71fae1732d1c · inbound

Natural Language Fine-Tuning cites this paper.

Natural Language Fine-Tuning DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:29:08.694519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:29:08.694519Z digest=sha256:99a45cb9b07a49cd538872ff3bb11276b5db8de973ee396e266a63295f43b9ab

Observation 63bb8363-8867-4ca0-aea9-6a5f3b3bc8d3 · inbound

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning cites this paper.

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T22:20:13.900296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:20:13.900296Z digest=sha256:30669f0d308b6a532e9757bb7164430d8280b4980892731fc860762550b53eb8

Observation 651c2a9e-9828-4a4a-bdb5-73a55ba1f21c · inbound

On Almost Surely Safe Alignment of Large Language Models at Inference-Time cites this paper.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.737239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.737239Z digest=sha256:5ade489a0eed0e1579ecbf47102e75bfff79132eac9c9493d749ed24ac2f06d7

Observation b252c8a5-7a7f-4a73-b37a-d5639dbc3f5a · inbound

Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms cites this paper.

Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T06:04:01.630684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T06:04:01.630684Z digest=sha256:e46a551670dda54fce18b4fb6c03a870391322bda7033d781b5dcf34580b33a8

Observation 0247b814-5dab-4ed2-8f03-997e2df12d94 · inbound

PIPA: Preference Alignment as Prior-Informed Statistical Estimation cites this paper.

PIPA: Preference Alignment as Prior-Informed Statistical Estimation DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:53.450848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:53.450848Z digest=sha256:0babc8de72f5802d73e5cc826c9c8b5d5efe7d3176a3fb1899532b0eb549cc79

Observation 1dc7d0db-9095-4dc4-bb4b-a1b3ed8c2c08 · inbound

Learning Explainable Dense Reward Shapes via Bayesian Optimization cites this paper.

Learning Explainable Dense Reward Shapes via Bayesian Optimization DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:13:53.739536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:13:53.739536Z digest=sha256:831c16e14ec215268571c7b4ee0fe03ae8dd037c3e3e95fb806ca618e3e86f6d

Observation 8a3e0111-a811-438a-bec5-fb8bb176282f · inbound

Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL cites this paper.

Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T00:59:21.917822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:59:21.917822Z digest=sha256:f1f18e80dd00b78d0c5a81bf06cf518c68b222b8ed4f16976742954d84cd78ea

Observation 31d2408d-5aec-4480-9b71-556bd5d2edc5 · inbound

A Survey on Progress in LLM Alignment from the Perspective of Reward Design cites this paper.

A Survey on Progress in LLM Alignment from the Perspective of Reward Design DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T00:52:06.794368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:52:06.794368Z digest=sha256:fbe3ef931408d3123c3f9a213eba267ec1936cfdcdd0f315a9751b2bc21a4453

Observation c2b93ec8-e704-4b73-ba47-f0276b1a35dc · inbound

Policy-labeled Preference Learning: Is Preference Enough for RLHF? cites this paper.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.865761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.865761Z digest=sha256:3ca641ee1adbf5442cd49baa3724ac5b6fa447f904dbf66ef3d5256675b5b7b1

Observation 52b3d24c-7bcc-479b-a14d-a47501771163 · inbound

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans cites this paper.

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:59.377307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:59.377307Z digest=sha256:2b4ac56ca50daa2c5cd68bc7a8c5da34b536d8095145b4ab3489ef9bdd1d08a7

Observation 961f0705-dff2-40ce-b452-174b38088b9e · inbound

EfficientLLM: Efficiency in Large Language Models cites this paper.

EfficientLLM: Efficiency in Large Language Models DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:35.120113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:35.120113Z digest=sha256:274ac9d1e0a3dd13c88305cef728a621df411f15c84e8cb6250724993b1281df

Observation 889509a3-c750-4706-85c5-c48ab90abe70 · inbound

Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning cites this paper.

Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:54.045722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:54.045722Z digest=sha256:778ce04c9716c27d1ac88fe58c7da186c516fc0362f1cfcc2fae5478ba382f2b

Observation d3bbb883-9740-4886-bf2b-91166959f7c4 · inbound

Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning cites this paper.

Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:36.509160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:36.509160Z digest=sha256:bb8dbc25e7c563820b32ba62c6f196a75a25c725b29d01c13d02d83cc7ae21b3

Observation c4de7e60-b185-4ffc-a5e9-0029ea61649a · inbound

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models cites this paper.

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:32.199965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:32.199965Z digest=sha256:9cc5166fbc830070d8e2fc2eddbd9836780b1b489828c75a27e6825447cf934f

Observation c31fd329-18ae-4c4a-af35-5ef3063ac485 · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.379167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.379167Z digest=sha256:a3b5ac87d58f38e34327f3bf5617bb0b26cdb8b452ed1dab62254dfcd0bc17f2

Observation 4dc8e471-1f47-4447-9f08-d0e01cf6afef · inbound

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training cites this paper.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.770706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.770706Z digest=sha256:29e3314af871e0762945793e83ad97685e29afae4b5c64b7487108c983c4a592

Observation 99c591d9-b88d-4fbb-b681-f880534ff6fc · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.297970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.297970Z digest=sha256:f8bda10bd2c0e08346b7e3fd5c38a1cec0f8fbc5f966acdfed2b1a1e7220c54d

Observation 4b519ca3-8bb9-4684-b640-aa5f040cee17 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 241

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.171442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.171442Z digest=sha256:9545b6dfef45dde8133bdf93004ecdaa5b86d210db784c0f27965334179705bc

Observation 50ee1f6f-2450-4e74-b775-ef523dadca71 · inbound

Reinforced Language Models for Sequential Decision Making cites this paper.

Reinforced Language Models for Sequential Decision Making DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:56.616695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:56.616695Z digest=sha256:57fee0ed6c003e6ee915e25c50557fc3f8aa449f3e5a2613b0a8e166decc92e7

Observation 49e5d7b9-3262-4791-be36-9cff586dd82b · inbound

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training cites this paper.

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:30:35.448387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T20:30:17.581842Z digest=sha256:a6580da23a520f5bb2c73feefcc28171b4f37963fe3a9880378980d6c775aa96

Observation 53e47d89-a31c-4e6e-af69-5298c870e248 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:45.181190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:45.181190Z digest=sha256:530643c25207e9e39a741495a01daa9325d900451c6abbbb21a7d7f9be7cf3e8

Observation 221083bd-c429-4d9a-9d64-47c01145b324 · inbound

Stabilizing Policy Optimization via Logits Convexity cites this paper.

Stabilizing Policy Optimization via Logits Convexity DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:09.115910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:09.115910Z digest=sha256:570300cbb44145f525cdb85298df68a81789bbb6bf6b7bd4084c7a9135ed0a6a

Observation c6edb85f-1d54-4e2e-b47c-4c321981fc4f · inbound

Data Agent: Learning to Select Data via End-to-End Dynamic Optimization cites this paper.

Data Agent: Learning to Select Data via End-to-End Dynamic Optimization DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:26:10.907667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T15:23:23.950152Z digest=sha256:c9d5f1ceac45699698d75b152f1944c6638129ec6c7f93b26534722242ae4a1c

Observation 67517349-28f0-43f7-bcc5-089ed583c2bc · inbound

LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO cites this paper.

LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:02.137107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T17:15:04.225059Z digest=sha256:50d4a9fa19fdadbc718fa84edace02e7f24a2d65edf6f8e63307e20ccbe08f50

Observation 4f94fb72-2b09-4a9c-bbe7-9d8a3c520e93 · inbound

Leveraging RAG for Training-Free Alignment of LLMs cites this paper.

Leveraging RAG for Training-Free Alignment of LLMs DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:42:08.175821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-13T02:41:07.016045Z digest=sha256:4d2312a4ab8bbc73447fd7baba059bb6328ac0cb1b50ac0d55e7301ce6b6ba92

Observation 5e8257b5-3f06-4443-a5d7-190b1c47cf5a · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:57:17.231849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:aa88e45e786a56a69079b8118f97826ce703834fab950f53c262ce4bd6a9566d

Observation 1d51c5b1-3850-4a7d-9b06-c6f90c87d07e · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:06.597442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:eaf319eaf2a4166cd078b5663b9629cefe18defa0cc8334638bda30956930962

Observation 3aaa981a-1563-4582-893a-02f8c5267fe0 · inbound

LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition cites this paper.

LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.011746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T18:44:14.878564Z digest=sha256:0eb5607582a6957a862dd428512ceaa926f8cc10d54d3e469bf1fdfac1985d7c

Observation bc0dafb7-0105-4645-91a7-0b0856f5d5c4 · inbound

Entropy Is Not Enough: Unlocking Effective Reinforcement Learning for Visual Reasoning via Vision-Anchored Token Selection cites this paper.

Entropy Is Not Enough: Unlocking Effective Reinforcement Learning for Visual Reasoning via Vision-Anchored Token Selection DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.893054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T09:58:34.557353Z digest=sha256:6ae664d3c290d0474f90f8129ac5bc416288961d6a3e4891679fbf10dfb9095f

Observation 697697b2-7bb0-4d44-972b-3bbb27b06c37 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 291

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:93b406850dbef3a11f8a2dda223201ad5ddee2b33cb1e84f0fdc5011e3f327eb

Observation 943e17f6-f6c4-493b-86dd-ea745a096419 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 292

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:06.696808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:06.696808Z digest=sha256:f0f092fba8d7a9db0abadb241cf273c908a7f873913202d55bce91323f480e01

Observation ee81150d-a934-48b1-8599-f78542a401ea · inbound

Distilled Reinforcement Learning for LLM Post-training cites this paper.

Distilled Reinforcement Learning for LLM Post-training DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:41.185696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:41.185696Z digest=sha256:c40ca43e10f22afd98c8ef8e6d7f1bade3e7641ea665c5c97a0d774625115bf1

Observation 2b2a4c63-8279-42a5-a4c3-9772c59ad7ec · inbound

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration cites this paper.

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 160

Resolution
unresolved
no resolver link, observed 2026-08-15T14:33:57.511242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:33:57.511242Z digest=sha256:2e7b9efc1b8f083560939e48654b5ea4d4a1b7075cc6939bea9ee58e0fc7ba7d

Observation 08fbb978-b2e7-4410-9e25-fa1a2c693315 · inbound

Token-Level Credit Assignment Optimization for Generative Document Retrieval cites this paper.

Token-Level Credit Assignment Optimization for Generative Document Retrieval DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:23:46.290938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:23:46.290938Z digest=sha256:73b997e76ee9ca83f323f3793ba8ba2b66dd31503f4f54a54a99e06e64629b8c