REVIEW 11 cited by
RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
read the original abstract
With the increasing use of large language models (LLMs), ensuring reliable performance in diverse, real-world environments is essential. Despite their remarkable achievements, LLMs often struggle with adversarial inputs, significantly impacting their effectiveness in practical applications. To systematically understand the robustness of LLMs, we present RUPBench, a comprehensive benchmark designed to evaluate LLM robustness across diverse reasoning tasks. Our benchmark incorporates 15 reasoning datasets, categorized into commonsense, arithmetic, logical, and knowledge-intensive reasoning, and introduces nine types of textual perturbations at lexical, syntactic, and semantic levels. By examining the performance of state-of-the-art LLMs such as GPT-4o, Llama3, Phi-3, and Gemma on both original and perturbed datasets, we provide a detailed analysis of their robustness and error patterns. Our findings highlight that larger models tend to exhibit greater robustness to perturbations. Additionally, common error types are identified through manual inspection, revealing specific challenges faced by LLMs in different reasoning contexts. This work provides insights into areas where LLMs need further improvement to handle diverse and noisy inputs effectively.
Forward citations
Cited by 11 Pith papers
-
Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam
A content-preserving red herring lowers LLM accuracy by 12.3 pp on sixty graduate microeconomics problems, corrupts reasoning without breaking coherence, and makes models rate the harder version as easier.
-
Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam
An irrelevant passage inserted into a graduate economics problem lowers LLM final-answer accuracy by 12.3 percentage points while models rate the corrupted problems as easier.
-
Understanding the Impact of Linguistic Realization Choices on LLM Stance with Causal Tracing
Controlled linguistic rewrites, including meaning-preserving ones, shift political stance judgments in open-weight LLMs, and activation patching localizes the restoration signal to mid-to-late final-position block outputs.
-
Implicit Reasoning Steering via Concept Chaining
Reinforcement-learning-optimized concept-chain paragraphs covertly steer language-model multiple-choice preferences after continued pretraining, with far lower detectability than direct paraphrases.
-
Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning
Seirênes trains LLMs via adversarial self-play to generate and overcome evolving distractions, producing gains of 7-10 points on math reasoning benchmarks and exposing blind spots in larger models.
-
GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations
GSM-SEM is a reusable framework for creating semantically variant augmentations of math benchmarks like GSM8K that alter facts but preserve answers and difficulty, with evaluations showing LLM performance drops of up ...
-
GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations
GSM-SEM generates reusable, stochastic semantic variants of math reasoning benchmarks that alter underlying facts but preserve answers, producing larger LLM performance drops than prior surface-level variants.
-
Towards a Science of AI Agent Reliability
Measuring 14 AI agents across two benchmarks, the paper finds 18 months of accuracy gains (≈0.21/yr) bought only small reliability gains (0.03–0.10/yr) under its consistency/robustness/predictability/safety framework.
-
Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam
Inserting an irrelevant passage into graduate microeconomics problems lowers LLM final-answer accuracy by 12.3 percentage points, corrupts the reasoning, preserves response form, and makes models rate the corrupted ta...
-
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
LLM reasoning benchmark scores vary substantially across repeated runs under the same model, strategy, and task, so single-run evaluation can misrank systems.
-
When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations
Evaluation shows LLMs for healthcare are sensitive to prompt changes, leading to inconsistent and potentially harmful clinical outputs on MedMCQA.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.