ToxiShield combines BERT for 97% F1 toxicity detection, Claude for multiclass classification, and Llama for detoxification achieving 95% style transfer accuracy on code review texts.
Title resolution pending
2 Pith papers cite this work, alongside 35 external citations. Polarity classification is still indexing.
2
Pith papers citing it
35
external citations · OpenAlex
citation-role summary
background 1
citation-polarity summary
fields
cs.SE 2years
2026 2verdicts
UNVERDICTED 2roles
background 1polarities
background 1representative citing papers
LLMs reach moderate macro-F1 scores of 0.36-0.37 when classifying code review comments into six smells and three useful intents, with one-shot examples helping some models on intent labels.
citing papers explorer
-
Real-Time Toxicity Filtering for Open-Source Code Reviews
ToxiShield combines BERT for 97% F1 toxicity detection, Claude for multiclass classification, and Llama for detoxification achieving 95% style transfer accuracy on code review texts.
-
Automated Classification of Human Code Review Comments with Large Language Models
LLMs reach moderate macro-F1 scores of 0.36-0.37 when classifying code review comments into six smells and three useful intents, with one-shot examples helping some models on intent labels.