GRAID augments small harmful-text datasets by generating embedding-guided synthetic examples and then filtering them through a multi-agentic LLM evaluation loop, improving downstream guardrail classifier F1 on BeaverTails and WildGuard.
Values show percentage of failed generations evaluated by � � � and � � � in each evaluation cycle
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Styleclone: Face Stylization with Diffusion Based Data Augmentation
GRAID augments small harmful-text datasets by generating embedding-guided synthetic examples and then filtering them through a multi-agentic LLM evaluation loop, improving downstream guardrail classifier F1 on BeaverTails and WildGuard.