Pith. sign in

A Survey of Direct Preference Optimization

7 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.

7 Pith papers citing it
2 external citations · external index

citation-role summary

background 2 method 1

citation-polarity summary

years

2026 6 2025 1

representative citing papers

TUX: Measuring Human--AI Tacit Understanding

cs.HC · 2026-05-29 · unverdicted · novelty 6.0

Profile-conditioned LLMs achieve higher tacit alignment with humans on subjective spectra when traits match, as quantified by the new Tacit Understanding Index (TUX) from 241 humans and 200 agents.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment

cs.CL · 2025-08-25 · unverdicted · novelty 5.0

Reinforced Behavior Alignment (RBA) uses self-synthesized data from a teacher LLM and reinforcement learning to close the instruction-following gap in SpeechLMs, outperforming distillation and reaching SOTA on spoken QA and speech-to-text translation benchmarks.

citing papers explorer

Showing 7 of 7 citing papers.