Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 39 inbound Pith citation observations for arXiv:2404.11999.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:43.539104Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T10:48:02.074671Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 078580d8-a1a5-4ca6-b1db-d8247483d942 · inbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Token-level Direct Preference Optimization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ba846ef6-2e07-4e29-a80a-bb825c1e5a78 · inbound
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Token-level Direct Preference Optimization
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 412575e6-f16c-4f16-b4a6-be6655f7d600 · inbound
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability Token-level Direct Preference Optimization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44648541-cb76-49be-8203-d9cce74a5e55 · inbound
T-REG: Preference Optimization with Token-Level Reward Regularization Token-level Direct Preference Optimization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df554fa1-4c4e-4e47-8839-3d33f48fea7a · inbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Token-level Direct Preference Optimization
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5614f044-9ccb-4a5a-902d-271b8ddd9c13 · inbound
SDPO: Segment-Level Direct Preference Optimization for Social Agents Token-level Direct Preference Optimization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 404700b7-9a2b-4d46-9a65-bf8dc02287f0 · inbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Token-level Direct Preference Optimization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10e4deab-3929-4688-bbd8-84862fb74085 · inbound
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Token-level Direct Preference Optimization
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cd4cba5-630c-4b7c-9577-6982d9e9e87a · inbound
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm Token-level Direct Preference Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd73f9b8-ffdf-4584-8b9a-66948967a9b3 · inbound
A Survey on Progress in LLM Alignment from the Perspective of Reward Design Token-level Direct Preference Optimization
Reference 106
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6c3cc89-936e-48cb-92b4-9bd3348890fb · inbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Token-level Direct Preference Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 113a36c2-1c1d-4bd1-90a2-8ee91dc66bb3 · inbound
SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment Token-level Direct Preference Optimization
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cc37db7-4a5c-40ce-99ae-84eaceba409b · inbound
Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Token-level Direct Preference Optimization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74b81ec1-c15b-4fb3-a0bb-fb078d10d9bd · inbound
Learning Safety Constraints for Large Language Models Token-level Direct Preference Optimization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afa98307-1de3-4953-b0fe-9cce2510cb6e · inbound
AI Agent Behavioral Science Token-level Direct Preference Optimization
Reference 186
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b868000e-f0a3-40b4-b6c6-1f723f5999fd · inbound
From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations Token-level Direct Preference Optimization
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c89f84c3-6e88-4ad9-9a0f-a6ab33c4fc8e · inbound
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap Token-level Direct Preference Optimization
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d8f6c5cd-fbb1-44ae-8344-3539981e24fd · inbound
FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus Token-level Direct Preference Optimization
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb7e1abf-09c4-4a9e-9392-c24a0c7a4d6c · inbound
Enhancing Speech Large Language Models through Reinforced Behavior Alignment Token-level Direct Preference Optimization
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 39310ff7-28d9-4fea-889c-7381e850af02 · inbound
Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning Token-level Direct Preference Optimization
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 107e3e7e-742e-4fc4-8ce9-5a2438b9cac2 · inbound
LLM Harms: A Taxonomy and Discussion Token-level Direct Preference Optimization
Reference 226
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7edaff37-1b18-4831-ae58-90fd720dcab3 · inbound
LLM Harms: A Taxonomy and Discussion Token-level Direct Preference Optimization
Reference 226
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e113064d-c96d-4e01-a4e1-845eef5436fa · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Token-level Direct Preference Optimization
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b989e5c4-d7dc-4c55-9f54-8ec5a7a0c8d1 · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Token-level Direct Preference Optimization
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1004e4b8-4187-433f-a6bd-8bc92db8614c · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning Token-level Direct Preference Optimization
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e8366876-ff4a-4a68-9242-88144359e01e · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning Token-level Direct Preference Optimization
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05ac2df2-b9ff-45cf-9199-d8e5fcbf58e0 · inbound
RAD-DPO: Robust Adaptive Denoising Direct Preference Optimization for Generative Retrieval in E-commerce Token-level Direct Preference Optimization
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 658ce455-d299-478c-9890-70c9f127eba5 · inbound
Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization Token-level Direct Preference Optimization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3fd45605-c623-4937-a472-9f913ea078d2 · inbound
Step-level Denoising-time Diffusion Alignment with Multiple Objectives Token-level Direct Preference Optimization
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1325efea-62bc-48c4-8707-c6fde731cd83 · inbound
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Token-level Direct Preference Optimization
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f0ca4aaf-90d5-4706-bdf6-b147accde66f · inbound
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Token-level Direct Preference Optimization
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 57c9d52e-3b80-44bf-9d55-2c43f27631dc · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Token-level Direct Preference Optimization
Reference 192
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 94827eba-847a-4d7b-9b89-3502ab126c78 · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Token-level Direct Preference Optimization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a8e2813e-a815-4476-8a2a-e98ae7d2016c · inbound
DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning Token-level Direct Preference Optimization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a5ba103b-130b-4b35-b5e4-2252069b86d8 · inbound
Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs Token-level Direct Preference Optimization
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ac02c51d-db5c-49a4-8a67-44dc879a8e09 · inbound
Flow Reasoning Models: Scaling Reasoning Through Iterative Self-Refinement Token-level Direct Preference Optimization
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e631884d-355e-47d0-9e7c-6c823cd3d4cd · inbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Token-level Direct Preference Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cb38f9a-977e-4430-8c10-a4f24d11d9dc · inbound
Test-Time Scaling via Error Localization Token-level Direct Preference Optimization
Reference 182
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54293756-9f1b-415b-aaa1-5be4700cd828 · inbound
Token-Level Credit Assignment Optimization for Generative Document Retrieval Token-level Direct Preference Optimization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.