Pith. sign in

REVIEW 10 cited by

Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.05534 v1 pith:F5ZUBTTB submitted 2024-06-08 cs.AI cs.CLcs.LG

classification cs.AIcs.CLcs.LG
keywords onlinealignmentchasingoptimizationpreferencecompetitioncontinualcross-domain
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Direct Preference Optimization (DPO) improves the alignment of large language models (LLMs) with human values by training directly on human preference datasets, eliminating the need for reward models. However, due to the presence of cross-domain human preferences, direct continual training can lead to catastrophic forgetting, limiting DPO's performance and efficiency. Inspired by intraspecific competition driving species evolution, we propose a Online Fast-Slow chasing DPO (OFS-DPO) for preference alignment, simulating competition through fast and slow chasing among models to facilitate rapid adaptation. Specifically, we first derive the regret upper bound for online learning, validating our motivation with a min-max optimization pattern. Based on this, we introduce two identical modules using Low-rank Adaptive (LoRA) with different optimization speeds to simulate intraspecific competition, and propose a new regularization term to guide their learning. To further mitigate catastrophic forgetting in cross-domain scenarios, we extend the OFS-DPO with LoRA modules combination strategy, resulting in the Cross domain Online Fast-Slow chasing DPO (COFS-DPO). This method leverages linear combinations of fast modules parameters from different task domains, fully utilizing historical information to achive continual value alignment. Experimental results show that OFS-DPO outperforms DPO in in-domain alignment, while COFS-DPO excels in cross-domain continual learning scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Test-Time Scaling via Error Localization

    cs.LG 2026-07 conditional novelty 6.0 of 10

    TTEL uses feedback-induced token probability drops to localize the first error in a failed reasoning trace and branch a new generation from that prefix, improving pass@k per token on coding and math benchmarks.

  2. Phi-Ground Tech Report: Advancing Perception in GUI Grounding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Phi-Ground models achieve state-of-the-art click accuracy on five GUI grounding benchmarks for models under 10B parameters using a 40M-sample training recipe with text-first inputs, random-resize augmentation, uniform...

  3. Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Applying DPO with preferences from people with intellectual disabilities improves German simplified-text readability but reduces semantic fidelity, and inconsistent target-group preferences prevent statistically signi...

  4. Do not Abstain! Identify and Solve the Uncertainty

    cs.AI 2025-06 conditional novelty 6.0 of 10

    The authors build ConfuseBench for three sources of LLM uncertainty and show that asking a follow-up question and checking whether its answer is unique improves source identification.

  5. Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator

    cs.AI 2025-02 conditional novelty 6.0 of 10

    Sampling dispreferred completions proportionally to the current reward model, as in contrastive divergence, improves preference-optimization performance and is framed as NLL estimation.

  6. Understanding the Logic of Direct Preference Alignment through Logic

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Direct preference alignment losses can be expressed as logical programs over model predictions, yielding an organized landscape of billions of definable losses and a route to new variants.

  7. Gene-R1: Reasoning with Data-Augmented Lightweight LLMs for Gene Set Analysis

    q-bio.GN 2025-09 conditional novelty 5.0 of 10

    Gene-R1 combines domain pre-training, GPT-o1 distilled reasoning examples, and GRPO reinforcement learning so lightweight Llama models match commercial LLMs on gene set analysis.

  8. DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management

    cs.LG 2025-05 conditional novelty 5.0 of 10

    DGRO decouples the KL regularization coefficient in reward optimization into two hyperparameters and shows strong reasoning results, though its reward-variance ablation is confounded.

  9. Direct Advantage Regression: Aligning LLMs with Online AI Reward

    cs.AI 2025-04 conditional novelty 5.0 of 10

    DAR aligns LLMs through advantage-weighted supervised fine-tuning on online AI scalar rewards with dual KL regularization, and reports win-rate gains over online RLHF and online preference methods.

  10. Salamandra Technical Report

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Salamandra is an open, from-scratch multilingual LLM family with 2B, 7B, and 40B checkpoints, instruction-tuned variants, a vision proof-of-concept, and detailed evaluations across Iberian and European languages.

Pith tools