Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 39 inbound Pith citation observations for arXiv:2404.18922.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:13:53.739536Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T03:26:29.891495Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 66dc5680-099b-439a-8e2a-4c969731c125 · inbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 948fde31-2f2d-4ef5-8106-19293324a369 · inbound
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96bb879e-6256-4140-988c-471195b71c1d · inbound
T-REG: Preference Optimization with Token-Level Reward Regularization DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4659949-2b58-4b5b-8532-ade787f3a012 · inbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cac1b905-4782-4861-851b-c2214bd75009 · inbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b2af11e-e5ae-4fd5-a7ee-71fae1732d1c · inbound
Natural Language Fine-Tuning DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63bb8363-8867-4ca0-aea9-6a5f3b3bc8d3 · inbound
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 651c2a9e-9828-4a4a-bdb5-73a55ba1f21c · inbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b252c8a5-7a7f-4a73-b37a-d5639dbc3f5a · inbound
Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0247b814-5dab-4ed2-8f03-997e2df12d94 · inbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dc7d0db-9095-4dc4-bb4b-a1b3ed8c2c08 · inbound
Learning Explainable Dense Reward Shapes via Bayesian Optimization DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a3e0111-a811-438a-bec5-fb8bb176282f · inbound
Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31d2408d-5aec-4480-9b71-556bd5d2edc5 · inbound
A Survey on Progress in LLM Alignment from the Perspective of Reward Design DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2b93ec8-e704-4b73-ba47-f0276b1a35dc · inbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52b3d24c-7bcc-479b-a14d-a47501771163 · inbound
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 961f0705-dff2-40ce-b452-174b38088b9e · inbound
EfficientLLM: Efficiency in Large Language Models DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 889509a3-c750-4706-85c5-c48ab90abe70 · inbound
Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3bbb883-9740-4886-bf2b-91166959f7c4 · inbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4de7e60-b185-4ffc-a5e9-0029ea61649a · inbound
Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c31fd329-18ae-4c4a-af35-5ef3063ac485 · inbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dc8e471-1f47-4447-9f08-d0e01cf6afef · inbound
VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99c591d9-b88d-4fbb-b681-f880534ff6fc · inbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 116
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b519ca3-8bb9-4684-b640-aa5f040cee17 · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 241
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50ee1f6f-2450-4e74-b775-ef523dadca71 · inbound
Reinforced Language Models for Sequential Decision Making DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e5d7b9-3262-4791-be36-9cff586dd82b · inbound
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 53e47d89-a31c-4e6e-af69-5298c870e248 · inbound
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 124
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 221083bd-c429-4d9a-9d64-47c01145b324 · inbound
Stabilizing Policy Optimization via Logits Convexity DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6edb85f-1d54-4e2e-b47c-4c321981fc4f · inbound
Data Agent: Learning to Select Data via End-to-End Dynamic Optimization DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 67517349-28f0-43f7-bcc5-089ed583c2bc · inbound
LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4f94fb72-2b09-4a9c-bbe7-9d8a3c520e93 · inbound
Leveraging RAG for Training-Free Alignment of LLMs DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5e8257b5-3f06-4443-a5d7-190b1c47cf5a · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1d51c5b1-3850-4a7d-9b06-c6f90c87d07e · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3aaa981a-1563-4582-893a-02f8c5267fe0 · inbound
LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bc0dafb7-0105-4645-91a7-0b0856f5d5c4 · inbound
Entropy Is Not Enough: Unlocking Effective Reinforcement Learning for Visual Reasoning via Vision-Anchored Token Selection DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 697697b2-7bb0-4d44-972b-3bbb27b06c37 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 291
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 943e17f6-f6c4-493b-86dd-ea745a096419 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 292
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee81150d-a934-48b1-8599-f78542a401ea · inbound
Distilled Reinforcement Learning for LLM Post-training DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b2a4c63-8279-42a5-a4c3-9772c59ad7ec · inbound
Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 160
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08fbb978-b2e7-4410-9e25-fa1a2c693315 · inbound
Token-Level Credit Assignment Optimization for Generative Document Retrieval DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.