A single-pass local LLM plus a random-forest validator flags likely annotation errors, letting two humans review 1.5% of 124,757 GitHub conversations and yielding a 946-toxic-conversation dataset that revises several small-scale findings about OSS toxicity.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Low-Cost Human-in-the-Loop Investigation of Toxicity on GitHub at Scale
A single-pass local LLM plus a random-forest validator flags likely annotation errors, letting two humans review 1.5% of 124,757 GitHub conversations and yielding a 946-toxic-conversation dataset that revises several small-scale findings about OSS toxicity.