A new 200-prompt benchmark with age-group splits and a 0-5 refusal scale finds that even top LLMs are unsafe on 5-28% of child-facing adversarial prompts, with open-weight models performing worst.
Available: https://arxiv.org/abs/2502.12552v1
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CY 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Safe-Child-LLM: A Developmental Benchmark for Evaluating LLM Safety in Child-LLM Interactions
A new 200-prompt benchmark with age-group splits and a 0-5 refusal scale finds that even top LLMs are unsafe on 5-28% of child-facing adversarial prompts, with open-weight models performing worst.