Pith. sign in

REVIEW 21 cited by

Token-level Direct Preference Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.11999 v5 pith:BRHDIKEQ submitted 2024-04-18 cs.CL cs.AI

classification cs.CLcs.AI
keywords tdpodivergencegenerationmethodstokenalignalignmentdirect
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Fine-tuning pre-trained Large Language Models (LLMs) is essential to align them with human values and intentions. This process often utilizes methods like pairwise comparisons and KL divergence against a reference LLM, focusing on the evaluation of full answers generated by the models. However, the generation of these responses occurs in a token level, following a sequential, auto-regressive fashion. In this paper, we introduce Token-level Direct Preference Optimization (TDPO), a novel approach to align LLMs with human preferences by optimizing policy at the token level. Unlike previous methods, which face challenges in divergence efficiency, TDPO incorporates forward KL divergence constraints for each token, improving alignment and diversity. Utilizing the Bradley-Terry model for a token-based reward system, TDPO enhances the regulation of KL divergence, while preserving simplicity without the need for explicit reward modeling. Experimental results across various text tasks demonstrate TDPO's superior performance in balancing alignment with generation diversity. Notably, fine-tuning with TDPO strikes a better balance than DPO in the controlled sentiment generation and single-turn dialogue datasets, and significantly improves the quality of generated responses compared to both DPO and PPO-based RLHF methods. Our code is open-sourced at https://github.com/Vance0124/Token-level-Direct-Preference-Optimization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Token-Level Credit Assignment Optimization for Generative Document Retrieval

    cs.IR 2026-08 conditional novelty 6.0 of 10

    TCA, a token-level credit assignment reinforcement learning framework, improves R@1 and MRR@10 over sequence-level RL baselines in generative document retrieval on MS MARCO and Natural Questions.

  2. Test-Time Scaling via Error Localization

    cs.LG 2026-07 conditional novelty 6.0 of 10

    TTEL uses feedback-induced token probability drops to localize the first error in a failed reasoning trace and branch a new generation from that prefix, improving pass@k per token on coding and math benchmarks.

  3. Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A token-level correctness classifier trained with LoRA and then merged into the model boosts out-of-distribution factuality in summarization and translation.

  4. rePIRL: Learn PRM with Inverse RL for LLM Reasoning

    cs.LG 2026-02 unverdicted novelty 6.0 of 10

    rePIRL learns token-level process rewards for LLM reasoning via a guided-cost-learning-style IRL objective, and shows these rewards improve reasoning policies on math/coding benchmarks.

  5. Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning

    cs.LG 2025-10 conditional novelty 6.0 of 10

    DiPO is a distribution-level unlearning method that constructs preference distributions from the model's own high-confidence logits and achieves state-of-the-art forget quality on TOFU while preserving utility.

  6. FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus

    cs.CV 2025-09 conditional novelty 6.0 of 10

    FocusDPO adds dynamic spatial weighting to preference-based fine-tuning, improving subject fidelity and reducing attribute leakage in multi-subject personalized image generation.

  7. Learning Safety Constraints for Large Language Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A polytope learned in LLM representation space can detect unsafe regions and steer outputs back to safety at inference time, reducing jailbreak success across several models.

  8. Probability-Consistent Preference Optimization for Enhanced LLM Reasoning

    cs.CL 2025-05 conditional novelty 6.0 of 10

    PCPO selects preference pairs by combining correct-answer status with token-level probability consistency, then trains with a weighted DPO+NLL loss, yielding small and partly inconsistent gains over outcome-only metho...

  9. SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment

    cs.LG 2025-05 conditional novelty 6.0 of 10

    SGDPO modifies DPO with a subsequence-based pilot term and reports up to 9.19% relative MT-Bench gains, though the claimed gradient mechanism is not fully derived for the implemented loss.

  10. Policy-labeled Preference Learning: Is Preference Enough for RLHF?

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Policy-labeled Preference Learning models preferences with regret and behavior-policy labels, adds contrastive KL regularization, and reports improved RLHF performance on MetaWorld offline and online control tasks.

  11. VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A hierarchical DPO training method and dataset reduce hallucination in video LLMs by aligning preferences at video, clip, object, and token levels.

  12. PIPA: Preference Alignment as Prior-Informed Statistical Estimation

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A unified maximum-likelihood framework with prior constraints that recovers DPO and KTO as special cases and yields new PIPA-M/PIPA-N losses with 3-10% gains on GSM8K and MATH.

  13. SDPO: Segment-Level Direct Preference Optimization for Social Agents

    cs.AI 2025-01 conditional novelty 6.0 of 10

    SDPO trains social agents by applying DPO-style preference optimization to equal-length key segments around an erroneous turn, and reports state-of-the-art SOTOPIA scores.

  14. T-REG: Preference Optimization with Token-Level Reward Regularization

    cs.CL 2024-12 conditional novelty 6.0 of 10

    T-REG adds contrastive-prompt self-generated token-level rewards as a weighted regularizer to DPO and SimPO, improving Alpaca Eval 2 length-controlled win rate by up to 3.8% and Arena-Hard win rate by up to 4.4% over ...

  15. Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability

    cs.CL 2024-11 conditional novelty 6.0 of 10

    A contrastive token-scoring method identifies 'critical tokens' in incorrect reasoning traces and penalizes them in DPO, yielding small but consistent accuracy gains on math benchmarks.

  16. Reward Modeling with Ordinal Feedback: Wisdom of the Crowd

    cs.LG 2024-11 conditional novelty 6.0 of 10

    The paper generalizes Bradley-Terry reward modeling to ordinal feedback labels and proves that, under a marginal unbiasedness assumption, finer-grained labels reduce Rademacher complexity and can improve reward learning.

  17. AI Agent Behavioral Science

    q-bio.NC 2025-06 conditional novelty 4.0 of 10

    AI agents should be studied as behavioral entities shaped by context and interaction, not only as trained models.

  18. A Survey on Progress in LLM Alignment from the Perspective of Reward Design

    cs.CL 2025-05 conditional novelty 4.0 of 10

    This paper organizes the LLM alignment literature into a reward-design-centered taxonomy and claims the field's evolution runs from rule-based to learned rewards and from RL-based to RL-free optimization.

  19. Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm

    cs.AI 2025-05 reject novelty 4.0 of 10

    Training 2D-DPO with an expected loss over uniform segment-score perturbations yields higher win rates under score noise than a clean-trained 2D-DPO baseline, though the comparison is confounded.

  20. Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

    cs.LG 2024-12 reject novelty 4.0 of 10

    Cal-DPO modifies DPO by adding squared losses that anchor the policy's implicit rewards of chosen and rejected responses to plus or minus 1/(2*beta), reporting gains on reasoning, summarization, and dialogue benchmarks.

  21. LLM Harms: A Taxonomy and Discussion

    cs.CY 2025-12 unverdicted novelty 3.0 of 10

    This paper proposes a taxonomy of LLM harms in five categories and suggests mitigation strategies plus a dynamic auditing system for responsible development.

Pith tools