Conditioning a multimodal guardrail on retrieved 'precedent' reasoning traces, rather than static policy definitions, improves few-shot and novel-policy content-moderation F1 scores on UnsafeBench.
Openai content policy
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Customize Multi-modal RAI Guardrails with Precedent-based predictions
Conditioning a multimodal guardrail on retrieved 'precedent' reasoning traces, rather than static policy definitions, improves few-shot and novel-policy content-moderation F1 scores on UnsafeBench.