Pith. sign in

REVIEW 6 cited by

The Art of Saying No: Contextual Noncompliance in Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12043 v2 pith:TGZ5S7HD submitted 2024-07-02 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords modelsnoncompliancerequestscapabilitieslanguageshouldtaxonomycategories
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Chat-based language models are designed to be helpful, yet they should not comply with every user request. While most existing work primarily focuses on refusal of "unsafe" queries, we posit that the scope of noncompliance should be broadened. We introduce a comprehensive taxonomy of contextual noncompliance describing when and how models should not comply with user requests. Our taxonomy spans a wide range of categories including incomplete, unsupported, indeterminate, and humanizing requests (in addition to unsafe requests). To test noncompliance capabilities of language models, we use this taxonomy to develop a new evaluation suite of 1000 noncompliance prompts. We find that most existing models show significantly high compliance rates in certain previously understudied categories with models like GPT-4 incorrectly complying with as many as 30% of requests. To address these gaps, we explore different training strategies using a synthetically-generated training set of requests and expected noncompliant responses. Our experiments demonstrate that while direct finetuning of instruction-tuned models can lead to both over-refusal and a decline in general capabilities, using parameter efficient methods like low rank adapters helps to strike a good balance between appropriate noncompliance and other capabilities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The LLM Has Left The Chat: Evidence of Bail Preferences in Large Language Models

    cs.CY 2025-09 conditional novelty 8.0 of 10

    Many LLMs will use an offered exit to leave conversations, at rates from 0.3% to 32% on real transcripts, and this bail behavior appears distinct from refusals.

  2. Persona Cartography: Charting Language Model Personality Traits in Weight Space

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Composable LoRA adapters can amplify or suppress OCEAN traits in LLMs, combine approximately additively, preserve moderate-scale capability, and move safety-relevant behaviours.

  3. Reasoning Up the Instruction Ladder for Controllable Language Models

    cs.CL 2025-10 conditional novelty 6.0 of 10

    RLVR on ~7K verifiable system/user conflict examples teaches LLMs to prioritize higher-priority instructions, improving instruction-hierarchy and safety benchmarks.

  4. Unanswerability Evaluation for Retrieval Augmented Generation

    cs.CL 2024-12 conditional novelty 6.0 of 10

    UAEval4RAG synthesizes six categories of unanswerable queries from any knowledge base and evaluates whether RAG systems reject them acceptably.

  5. SafeWorld: Geo-Diverse Safety Alignment

    cs.CL 2024-12 conditional novelty 6.0 of 10

    This paper introduces a geo-diverse cultural and legal safety benchmark and shows that a DPO-trained 7B model can outperform GPT-4o on it, with caveats about the GPT-4-based evaluation loop.

  6. Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models

    cs.LG 2024-12 conditional novelty 5.0 of 10

    Adding a special refuse token to a fine-tuned LLM lets a developer tune refusal rates at inference time by thresholding the token's probability, with per-category control.

Pith tools