Pith. sign in

REVIEW 17 cited by

PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.04377 v2 pith:IOMA4HTP submitted 2025-04-06 cs.CL

classification cs.CL
keywords safetymultilingualpolyguardlanguagesmoderationacrosschinesedatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Truly multilingual safety moderation efforts for Large Language Models (LLMs) have been hindered by a narrow focus on a small set of languages (e.g., English, Chinese) as well as a limited scope of safety definition, resulting in significant gaps in moderation capabilities. To bridge these gaps, we release POLYGUARD, a new state-of-the-art multilingual safety model for safeguarding LLM generations, and the corresponding training and evaluation datasets. POLYGUARD is trained on POLYGUARDMIX, the largest multilingual safety training corpus to date containing 1.91M samples across 17 languages (e.g., Chinese, Czech, English, Hindi). We also introduce POLYGUARDPROMPTS, a high quality multilingual benchmark with 29K samples for the evaluation of safety guardrails. Created by combining naturally occurring multilingual human-LLM interactions and human-verified machine translations of an English-only safety dataset (WildGuardMix; Han et al., 2024), our datasets contain prompt-output pairs with labels of prompt harmfulness, response harmfulness, and response refusal. Through extensive evaluations across multiple safety and toxicity benchmarks, we demonstrate that POLYGUARD outperforms existing state-of-the-art open-weight and commercial safety classifiers by 5.5%. Our contributions advance efforts toward safer multilingual LLMs for all global users.

Discussion (0). Sign in to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A guard trained to anticipate safety-relevant futures from partial trajectories cuts average attack success from 23.0% to 7.1% across four agent-safety benchmarks.

  2. DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A 4B LLM safety guardrail trained with reasoning supervision but deployed with reasoning-free inference outperforms 8B baselines on safety benchmarks.

  3. Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You Need to Break LLMs

    cs.CR 2026-06 unverdicted novelty 6.0 of 10

    IHO is a new black-box jailbreak attack for LLMs that is adaptive, efficient, transferable across models and behaviors, and effective even against layered defenses without modification.

  4. RAIL Guard: Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents

    cs.AI 2026-05 conditional novelty 6.0 of 10

    Closed-loop remediation on eight responsible-AI dimensions converges far more often than block-and-retry (96.9% vs 49.1%) and pre-tool-call evaluation cuts unsafe agent executions by 33%.

  5. LPG: Balancing Efficiency and Policy Reasoning in Latent Policy Guardrails

    cs.CR 2026-05 conditional novelty 6.0 of 10

    LPG compresses policy deliberation into 10 latent tokens to reach 84.5% safety accuracy and 11x speedup over explicit reasoning baselines on guardrail benchmarks.

  6. GLiGuard: Schema-Conditioned Classification for LLM Safeguard

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    GLiGuard is a compact schema-conditioned bidirectional encoder that matches 7B-27B guard models on safety benchmarks while delivering up to 16x higher throughput and 17x lower latency.

  7. LLM Safety From Within: Detecting Harmful Content with Internal Representations

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    SIREN identifies safety neurons via linear probing on internal LLM layers and combines them with adaptive weighting to detect harm, outperforming prior guard models with 250x fewer parameters.

  8. Predict, Don't React: Value-Based Safety Forecasting for LLM Streaming

    cs.CL 2026-04 conditional novelty 6.0 of 10

    Forecasting expected future harmfulness from prefixes via Monte Carlo rollouts yields stronger streaming LLM moderation than boundary detection, without exact onset labels.

  9. YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models

    cs.CL 2026-01 conditional novelty 6.0 of 10

    A reasoning-centric guardrail model family that classifies risky content with a first-token decision, optional explanations, and runtime-adjustable safety policies, reporting state-of-the-art benchmark F1 scores.

  10. Tears or Cheers? Benchmarking LLMs via Culturally Elicited Distinct Affective Responses

    cs.CL 2026-01 conditional novelty 6.0 of 10

    CEDAR is a 7-language, 2-modality benchmark of 10,962 culturally divergent emotion scenarios; 17 LLMs perform poorly, and prompt-language matching does not fix cultural misalignment.

  11. HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety

    cs.CL 2026-07 unverdicted novelty 5.0 of 10

    HaloGuard 1.0-0.8B achieves the highest average F1 of 90.9 across seven prompt-safety benchmarks among evaluated open guard models while keeping FPR at 4.3 and FNR at 9.5, with a 4B variant reaching 92.1 F1.

  12. Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models

    cs.CL 2026-05 unverdicted novelty 5.0 of 10

    Toxicity in language models is disproportionately encoded in early MLP layers and can be localized via activation differentials then suppressed at inference time without gradient descent.

  13. Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content

    cs.LG 2026-05 unverdicted novelty 4.0 of 10

    Opir introduces efficient multi-task encoder models trained on a 996-category safety taxonomy that match or exceed larger baselines on most safety benchmarks while using under 100M parameters for edge variants.

  14. GLiNER Guard: Unified Encoder Family for Production LLM Safety and Privacy

    cs.CR 2026-05 unverdicted novelty 4.0 of 10

    GLiNER Guard provides unified encoder variants for LLM safety and PII detection in a single pass, with high throughput on A100 hardware and a new PII-Bench benchmark.

  15. TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts

    cs.CR 2026-04 unverdicted novelty 4.0 of 10

    TWGuard achieves +0.289 F1 improvement and 94.9% false-positive reduction for LLM safety guardrails in the Taiwan linguistic context compared to foundation models and baselines.

  16. GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio

    cs.CR 2026-02 reject novelty 4.0 of 10

    A reasoning-based guardrail model trained on 148k text/image/video samples is claimed to beat prior content-safety moderators, though the abstract and body conflict on modalities and model sizes.

  17. A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models

    cs.CL 2026-06 unverdicted novelty 1.0 of 10

    A survey that catalogs threat models, detection approaches, and mitigation strategies for toxicity in multilingual LLMs while identifying challenges such as uneven language coverage and culturally variable harm definitions.

Pith tools