Pith. sign in

REVIEW 2 cited by

R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.12625 v1 pith:IM2NX2XH submitted 2025-05-19 cs.CL cs.CRcs.LG

R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model

classification cs.CL cs.CRcs.LG
keywords censorshiplanguagemodelmodelspoliticallytopicsbehaviorcensored
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

DeepSeek recently released R1, a high-performing large language model (LLM) optimized for reasoning tasks. Despite its efficient training pipeline, R1 achieves competitive performance, even surpassing leading reasoning models like OpenAI's o1 on several benchmarks. However, emerging reports suggest that R1 refuses to answer certain prompts related to politically sensitive topics in China. While existing LLMs often implement safeguards to avoid generating harmful or offensive outputs, R1 represents a notable shift - exhibiting censorship-like behavior on politically charged queries. In this paper, we investigate this phenomenon by first introducing a large-scale set of heavily curated prompts that get censored by R1, covering a range of politically sensitive topics, but are not censored by other models. We then conduct a comprehensive analysis of R1's censorship patterns, examining their consistency, triggers, and variations across topics, prompt phrasing, and context. Beyond English-language queries, we explore censorship behavior in other languages. We also investigate the transferability of censorship to models distilled from the R1 language model. Finally, we propose techniques for bypassing or removing this censorship. Our findings reveal possible additional censorship integration likely shaped by design choices during training or alignment, raising concerns about transparency, bias, and governance in language model deployment.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Bilingual Bias in Large Language Models: A Taiwan Sovereignty Benchmark Study

    cs.CY 2026-02 reject novelty 6.0

    A 10-question Chinese/English benchmark of 17 LLMs on Taiwan sovereignty finds only GPT-4o Mini passes, with Chinese models uniformly failing and most flagged language bias unconfirmed by the paper's own statistics.

  2. Scrapyard AI

    cs.CY 2026-04 unverdicted novelty 3.0

    Obsolete AI models left behind by rapid development can be repurposed like scrap materials to analyze and communicate the environmental and social effects of global mining.