Pith. sign in

REVIEW 4 cited by

A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.09431 v1 pith:VYY57AHZ submitted 2025-01-16 cs.AI cs.CLcs.CRcs.CY

classification cs.AIcs.CLcs.CRcs.CY
keywords llmssurveyalignmentapplicationscomprehensiveenhancinginherentprivacy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While large language models (LLMs) present significant potential for supporting numerous real-world applications and delivering positive social impacts, they still face significant challenges in terms of the inherent risk of privacy leakage, hallucinated outputs, and value misalignment, and can be maliciously used for generating toxic content and unethical purposes after been jailbroken. Therefore, in this survey, we present a comprehensive review of recent advancements aimed at mitigating these issues, organized across the four phases of LLM development and usage: data collecting and pre-training, fine-tuning and alignment, prompting and reasoning, and post-processing and auditing. We elaborate on the recent advances for enhancing the performance of LLMs in terms of privacy protection, hallucination reduction, value alignment, toxicity elimination, and jailbreak defenses. In contrast to previous surveys that focus on a single dimension of responsible LLMs, this survey presents a unified framework that encompasses these diverse dimensions, providing a comprehensive view of enhancing LLMs to better serve real-world applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Localizing Persona Representations in LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Persona information is most separable in the final third of LLM layers, and in Llama3's last layer ethical personas share 17.6% of salient activations while political personas have 2.1% to 5.5% unique activations.

  2. OrgAccess: A Benchmark for Role Based Access Control in Organization Scale LLMs

    cs.AI 2025-05 conditional novelty 6.0 of 10

    OrgAccess, a 70k-query synthetic RBAC benchmark, shows current LLMs including GPT-4.1 (F1 0.27 on the hardest split) struggle badly with permission adherence.

  3. AgentStealth: Reinforcing Large Language Model for Anonymizing User-generated Text

    cs.CL 2025-06 conditional novelty 5.0 of 10

    An 8B local language model, trained with an adversarial self-play workflow and reinforcement learning, anonymizes user-generated text with 12.3% better privacy protection and 6.8% better utility than prior LLM-based a...

  4. A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A position paper arguing that LLM-based human-agent systems, not fully autonomous agents, should be the immediate goal for AI development.

Pith tools