REVIEW 5 cited by
A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While large language models (LLMs) present significant potential for supporting numerous real-world applications and delivering positive social impacts, they still face significant challenges in terms of the inherent risk of privacy leakage, hallucinated outputs, and value misalignment, and can be maliciously used for generating toxic content and unethical purposes after been jailbroken. Therefore, in this survey, we present a comprehensive review of recent advancements aimed at mitigating these issues, organized across the four phases of LLM development and usage: data collecting and pre-training, fine-tuning and alignment, prompting and reasoning, and post-processing and auditing. We elaborate on the recent advances for enhancing the performance of LLMs in terms of privacy protection, hallucination reduction, value alignment, toxicity elimination, and jailbreak defenses. In contrast to previous surveys that focus on a single dimension of responsible LLMs, this survey presents a unified framework that encompasses these diverse dimensions, providing a comprehensive view of enhancing LLMs to better serve real-world applications.
Forward citations
Cited by 5 Pith papers
-
Localizing Persona Representations in LLMs
Persona information is most separable in the final third of LLM layers, and in Llama3's last layer ethical personas share 17.6% of salient activations while political personas have 2.1% to 5.5% unique activations.
-
OrgAccess: A Benchmark for Role Based Access Control in Organization Scale LLMs
OrgAccess, a 70k-query synthetic RBAC benchmark, shows current LLMs including GPT-4.1 (F1 0.27 on the hardest split) struggle badly with permission adherence.
-
Membership Inference Risks in Quantized Models: A Theoretical and Empirical Study
Quantizers can be ranked by privacy using r_Q, a rate constant built from the loss gap and variance of low-loss quantized checkpoints along the training trajectory.
-
AgentStealth: Reinforcing Large Language Model for Anonymizing User-generated Text
An 8B local language model, trained with an adversarial self-play workflow and reinforcement learning, anonymizes user-generated text with 12.3% better privacy protection and 6.8% better utility than prior LLM-based a...
-
A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy
A position paper arguing that LLM-based human-agent systems, not fully autonomous agents, should be the immediate goal for AI development.
Discussion (0). Continue with ORCID to comment.