Pith. sign in

emnlp-main.882/

7 Pith papers cite this work, alongside 32 external citations. Polarity classification is still indexing.

7 Pith papers citing it
32 external citations · OpenAlex

citation-role summary

background 1

citation-polarity summary

years

2026 7

roles

background 1

polarities

background 1

representative citing papers

Steerable Cultural Preference Optimization of Reward Models

cs.CL · 2026-06-17 · unverdicted · novelty 5.0

SCPO is a steerable training method for reward models that improves minority cultural preference accuracy by up to 7 points and is up to 280% more data-efficient than standard finetuning on PRISM and GlobalOpinionQA datasets.

citing papers explorer

Showing 7 of 7 citing papers.