REVIEW 10 cited by
Can LLMs Follow Simple Rules?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
As Large Language Models (LLMs) are deployed with increasing real-world responsibilities, it is important to be able to specify and constrain the behavior of these systems in a reliable manner. Model developers may wish to set explicit rules for the model, such as "do not generate abusive content", but these may be circumvented by jailbreaking techniques. Existing evaluations of adversarial attacks and defenses on LLMs generally require either expensive manual review or unreliable heuristic checks. To address this issue, we propose Rule-following Language Evaluation Scenarios (RuLES), a programmatic framework for measuring rule-following ability in LLMs. RuLES consists of 14 simple text scenarios in which the model is instructed to obey various rules while interacting with the user. Each scenario has a programmatic evaluation function to determine whether the model has broken any rules in a conversation. Our evaluations of proprietary and open models show that almost all current models struggle to follow scenario rules, even on straightforward test cases. We also demonstrate that simple optimization attacks suffice to significantly increase failure rates on test cases. We conclude by exploring two potential avenues for improvement: test-time steering and supervised fine-tuning.
Forward citations
Cited by 10 Pith papers
-
Evaluating Language Model Reasoning about Confidential Information
PasswordEval shows frontier models frequently leak passwords or confidential information, jailbreaks worsen failures, and reasoning traces leak secrets even when final answers do not.
-
LLMs for Customized Marketing Content Generation and Evaluation at Scale
MarketingFM generates e-commerce ad copy with RAG and an LLM; AutoEval uses LLM-as-a-Judge plus rule checks and self-refines its prompts, with online tests showing significant clicks and impressions lifts but no signi...
-
Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
Placing demographic audience information in system prompts rather than user prompts shifts sentiment and ranking outputs across six commercial LLMs, but the design confounds position with instruction content.
-
EnSToM: Enhancing Dialogue Systems with Entropy-Scaled Steering Vectors for Topic Maintenance
EnSToM scales activation steering by layer-wise entropy, improving distractor refusal in task-oriented dialogues while preserving on-topic responses.
-
Security Steerability is All You Need
An LLM's ability to follow application-specific system-prompt guardrails (security steerability) is nearly uncorrelated with its resistance to standard jailbreak attacks, based on a new 240-case benchmark across 18 op...
-
IHEval: Evaluating Language Models on Following the Instruction Hierarchy
IHEval shows that current language models often follow lower-priority instructions over system messages, and simple prompting does not fix the problem.
-
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
RuleArena evaluates LLMs on realistic rule-guided reasoning and finds that even o1-preview solves only about half of the easiest problems and near zero of the hardest.
-
Engineering Trustworthy Agentic AI for Critical Systems
A survey claiming that agentic AI trustworthiness is a single cross-domain problem and outlining a framework for graded, certifiable assurance.
-
Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution
Prompt injection can make LLM agents leak personal data they observed while executing tasks, with measured attack success rates around 15-20 percent and password leakage much rarer.
-
WALL-E 2.0: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agents
WALL-E 2.0 improves LLM agents by encoding learned environment rules as executable code that corrects an LLM world model, lifting ALFWorld success to 98%.
Discussion (0). Continue with ORCID to comment.