Pith. sign in

REVIEW 7 cited by

Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.15092 v1 pith:P6NJZYXO submitted 2025-03-19 cs.CR cs.AIcs.CL

classification cs.CRcs.AIcs.CL
keywords modelssafetyevaluationdeepseekcontentlargecapabilitiesfindings
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This study presents the first comprehensive safety evaluation of the DeepSeek models, focusing on evaluating the safety risks associated with their generated content. Our evaluation encompasses DeepSeek's latest generation of large language models, multimodal large language models, and text-to-image models, systematically examining their performance regarding unsafe content generation. Notably, we developed a bilingual (Chinese-English) safety evaluation dataset tailored to Chinese sociocultural contexts, enabling a more thorough evaluation of the safety capabilities of Chinese-developed models. Experimental results indicate that despite their strong general capabilities, DeepSeek models exhibit significant safety vulnerabilities across multiple risk dimensions, including algorithmic discrimination and sexual content. These findings provide crucial insights for understanding and improving the safety of large foundation models. Our code is available at https://github.com/NY1024/DeepSeek-Safety-Eval.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Mask-GCG uses learnable masks to prune a minority of low-impact tokens from GCG attack suffixes, slightly improving speed while showing most tokens are necessary.

  2. VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    VisCRA jailbreaks multimodal LLMs by masking the most harmful image region and using a two-stage reasoning prompt to make the model infer and then comply.

  3. Practical Reasoning Interruption Attacks on Reasoning Large Language Models

    cs.CR 2025-05 conditional novelty 6.0 of 10

    A tiny prompt can force DeepSeek-R1's reasoning content to overflow into the final answer, yielding a practical denial-of-service attack and a new jailbreak route.

  4. Manipulating Multimodal Agents via Cross-Modal Prompt Injection

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A coordinated attack that embeds malicious cues in both visual and textual inputs can hijack black-box multimodal agents, outperforming single-modality prompt injection attacks.

  5. POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models

    cs.CR 2025-05 conditional novelty 5.0 of 10

    POISONCRAFT injects adversarial documents into a RAG knowledge base that are likely to be retrieved for arbitrary user queries and that steer the language model into recommending a fake URL.

  6. A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

    cs.CL 2025-09 conditional novelty 4.0 of 10

    A structured literature survey concluding that reasoning capabilities do not automatically make LLMs more trustworthy and can introduce new vulnerabilities in safety, robustness, and privacy.

  7. PRJ: Perception-Retrieval-Judgement for Generated Images

    cs.CV 2025-06 reject novelty 4.0 of 10

    A new safety checker for AI-generated images, built from a vision-language model, retrieval-augmented knowledge lookup, and an LLM judge, reports higher detection rates and category-level toxicity scores than three ex...

Pith tools