Introduces TBPO, which derives a Bregman-divergence density-ratio matching objective for token-level preference optimization that generalizes DPO while preserving the induced optimal policy.
arXiv preprint arXiv:2406.16235 , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
Synthesizes literature into a four-stage lifecycle framework for cyberbullying governance from detection through proactive intervention.
citing papers explorer
-
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching
Introduces TBPO, which derives a Bregman-divergence density-ratio matching objective for token-level preference optimization that generalizes DPO while preserving the induced optimal policy.
-
Cyberbullying Governance on Social Media: A Unified Framework from Content Identification to Intervention
Synthesizes literature into a four-stage lifecycle framework for cyberbullying governance from detection through proactive intervention.