African content moderators suffer high psychological distress from systemic labor conditions, with former moderators showing lasting impacts and corporate wellness programs proving ineffective.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 1
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
RLHF-aligned language models show increasing resistance to red teaming with scale up to 52B parameters, unlike prompted or rejection-sampled models, supported by a released dataset of 38,961 attacks.
citing papers explorer
-
Beyond Content Exposure: Systemic Factors Driving Moderators' Mental Health Crisis in Africa
African content moderators suffer high psychological distress from systemic labor conditions, with former moderators showing lasting impacts and corporate wellness programs proving ineffective.
-
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
RLHF-aligned language models show increasing resistance to red teaming with scale up to 52B parameters, unlike prompted or rejection-sampled models, supported by a released dataset of 38,961 attacks.