Sequential DPO-style training makes coalition policies reconstructable by arithmetic on singleton models, enabling linear-cost approximation of Shapley data values.
Title resolution pending
1 Pith paper cite this work, alongside 8 external citations. Polarity classification is still indexing.
1
Pith paper citing it
8
external citations · OpenAlex
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Shapley-based Data Valuation for LLM Alignment via Sequential Preference Optimization
Sequential DPO-style training makes coalition policies reconstructable by arithmetic on singleton models, enabling linear-cost approximation of Shapley data values.