Step-level human preference data collected via SimPref, then SFT+DPO, improves long-horizon social-simulation behavior of open-weight LLM agents on held-out events.
Proceedings of the ACM on Human-Computer In- teraction 9(2), 1–27 (May 2025)
1 Pith paper cite this work, alongside 8 external citations. Polarity classification is still indexing.
1
Pith paper citing it
8
external citations · OpenAlex
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Step-Level Preference Learning for Generative Agents in Social Simulations
Step-level human preference data collected via SimPref, then SFT+DPO, improves long-horizon social-simulation behavior of open-weight LLM agents on held-out events.