Pith. sign in

REVIEW 3 cited by

Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.10328 v2 pith:662WQYIH submitted 2021-06-18 cs.CL cs.CY

classification cs.CLcs.CY
keywords languagemetricsmodelmodelspalmsprocessbehaviordataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language models can generate harmful and biased outputs and exhibit undesirable behavior according to a given cultural context. We propose a Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets, an iterative process to significantly change model behavior by crafting and fine-tuning on a dataset that reflects a predetermined set of target values. We evaluate our process using three metrics: quantitative metrics with human evaluations that score output adherence to a target value, toxicity scoring on outputs; and qualitative metrics analyzing the most common word associated with a given social category. Through each iteration, we add additional training dataset examples based on observed shortcomings from evaluations. PALMS performs significantly better on all metrics compared to baseline and control models for a broad range of GPT-3 language model sizes without compromising capability integrity. We find that the effectiveness of PALMS increases with model size. We show that significantly adjusting language model behavior is feasible with a small, hand-curated dataset.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Moderating Harm: Benchmarking Large Language Models for Cyberbullying Detection in YouTube Comments

    cs.CL 2025-05 conditional novelty 5.0 of 10

    In a zero-shot benchmark on 5,080 YouTube comments, GPT-4.1 achieved the best F1 balance at 0.863, while Gemini favored recall and Claude favored precision.

  2. Governance-as-a-Service: A Multi-Agent Framework for AI System Compliance and Policy Enforcement

    cs.LG 2025-08 reject novelty 4.0 of 10

    An external policy-enforcement layer with a trust score claims to block risky AI-agent actions in simulations, but its core formula is inconsistent across the paper.

  3. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools