REVIEW 8 cited by
Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Although Large Language Models (LLMs) have demonstrated strong ability, they are further supposed to be controlled and guided by in real-world scenarios to be safe, accurate, and intelligent. This demands the possession of capability of LLMs. However, no prior work has made a clear evaluation of the inferential rule-following capability of LLMs. Previous studies that try to evaluate the inferential rule-following capability of LLMs fail to distinguish the inferential rule-following scenarios from the instruction-following scenarios. Therefore, this paper first clarifies the concept of inferential rule-following and proposes a comprehensive benchmark, RuleBench, to evaluate a diversified range of inferential rule-following abilities. Our experimental results on a variety of LLMs show that they are still limited in following rules. Our analysis based on the evaluation results provides insights into the improvements for LLMs toward a better inferential rule-following intelligent agent. We further propose Inferential Rule-Following Tuning (IRFT). The experimental results show that through IRFT, LLMs can learn abstract rule-following abilities from purely synthetic data and then generalize to RuleBench. The data and code can be found at: https://anonymous.4open.science/r/llm-rule-following-B3E3/
Forward citations
Cited by 8 Pith papers
-
Evaluating Language Model Reasoning about Confidential Information
PasswordEval shows frontier models frequently leak passwords or confidential information, jailbreaks worsen failures, and reasoning traces leak secrets even when final answers do not.
-
GENIE-ASI: Generative Instruction and Executable Code for Analog Subcircuit Identification
A training-free LLM pipeline can generate executable Python code that identifies analog subcircuits in flattened SPICE netlists, matching rule-based labels on simple blocks and partially on complex ones.
-
LLMs for Customized Marketing Content Generation and Evaluation at Scale
MarketingFM generates e-commerce ad copy with RAG and an LLM; AutoEval uses LLM-as-a-Judge plus rule checks and self-refines its prompts, with online tests showing significant clicks and impressions lifts but no signi...
-
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models
MedGUIDE tests whether LLMs follow structured NCCN cancer-care decision trees and finds that even medical LLMs often lag general models on this task.
-
Shuttle Between the Instructions and the Parameters of Large Language Models
SHIP jointly trains an encoder and decoder so a language model's soft-prompt parameters can be reconstructed from instructions and vice versa, improving instruction induction and inductive reasoning.
-
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
RuleArena evaluates LLMs on realistic rule-guided reasoning and finds that even o1-preview solves only about half of the easiest problems and near zero of the hardest.
-
Improve Rule Retrieval and Reasoning with Self-Induction and Relevance ReEstimate
Using an LLM to induce an abstract rule from a query and then re-ranking retrieved rules with an LLM prompt improves rule retrieval and downstream reasoning in most tested configurations.
-
Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
A survey that frames the core problem of LLM evaluation as 'evaluation generalization': finite test sets cannot scale with unbounded model capabilities, and proposes two transitions in evaluation design.
Discussion (0). Continue with ORCID to comment.