Pith. sign in

REVIEW 11 cited by

RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11020 v1 pith:L5LV6M5C submitted 2024-06-16 cs.CL

RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models

classification cs.CL
keywords llmsreasoningrobustnessdiversemodelsperturbationsbenchmarkdatasets
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

With the increasing use of large language models (LLMs), ensuring reliable performance in diverse, real-world environments is essential. Despite their remarkable achievements, LLMs often struggle with adversarial inputs, significantly impacting their effectiveness in practical applications. To systematically understand the robustness of LLMs, we present RUPBench, a comprehensive benchmark designed to evaluate LLM robustness across diverse reasoning tasks. Our benchmark incorporates 15 reasoning datasets, categorized into commonsense, arithmetic, logical, and knowledge-intensive reasoning, and introduces nine types of textual perturbations at lexical, syntactic, and semantic levels. By examining the performance of state-of-the-art LLMs such as GPT-4o, Llama3, Phi-3, and Gemma on both original and perturbed datasets, we provide a detailed analysis of their robustness and error patterns. Our findings highlight that larger models tend to exhibit greater robustness to perturbations. Additionally, common error types are identified through manual inspection, revealing specific challenges faced by LLMs in different reasoning contexts. This work provides insights into areas where LLMs need further improvement to handle diverse and noisy inputs effectively.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam

    econ.GN 2026-07 accept novelty 6.5

    A content-preserving red herring lowers LLM accuracy by 12.3 pp on sixty graduate microeconomics problems, corrupts reasoning without breaking coherence, and makes models rate the harder version as easier.

  2. Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam

    econ.GN 2026-07 conditional novelty 6.0

    An irrelevant passage inserted into a graduate economics problem lowers LLM final-answer accuracy by 12.3 percentage points while models rate the corrupted problems as easier.

  3. Understanding the Impact of Linguistic Realization Choices on LLM Stance with Causal Tracing

    cs.CL 2026-07 conditional novelty 6.0

    Controlled linguistic rewrites, including meaning-preserving ones, shift political stance judgments in open-weight LLMs, and activation patching localizes the restoration signal to mid-to-late final-position block outputs.

  4. Implicit Reasoning Steering via Concept Chaining

    cs.CL 2026-07 conditional novelty 6.0

    Reinforcement-learning-optimized concept-chain paragraphs covertly steer language-model multiple-choice preferences after continued pretraining, with far lower detectability than direct paraphrases.

  5. Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning

    cs.AI 2026-05 unverdicted novelty 6.0

    Seirênes trains LLMs via adversarial self-play to generate and overcome evolving distractions, producing gains of 7-10 points on math reasoning benchmarks and exposing blind spots in larger models.

  6. GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations

    cs.CL 2026-05 unverdicted novelty 6.0

    GSM-SEM is a reusable framework for creating semantically variant augmentations of math benchmarks like GSM8K that alter facts but preserve answers and difficulty, with evaluations showing LLM performance drops of up ...

  7. GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations

    cs.CL 2026-05 unverdicted novelty 6.0

    GSM-SEM generates reusable, stochastic semantic variants of math reasoning benchmarks that alter underlying facts but preserve answers, producing larger LLM performance drops than prior surface-level variants.

  8. Towards a Science of AI Agent Reliability

    cs.AI 2026-02 conditional novelty 6.0

    Measuring 14 AI agents across two benchmarks, the paper finds 18 months of accuracy gains (≈0.21/yr) bought only small reliability gains (0.03–0.10/yr) under its consistency/robustness/predictability/safety framework.

  9. Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam

    econ.GN 2026-07 conditional novelty 5.0

    Inserting an irrelevant passage into graduate microeconomics problems lowers LLM final-answer accuracy by 12.3 percentage points, corrupts the reasoning, preserves response form, and makes models rate the corrupted ta...

  10. ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning

    cs.AI 2025-12 reject novelty 5.0

    LLM reasoning benchmark scores vary substantially across repeated runs under the same model, strategy, and task, so single-run evaluation can misrank systems.

  11. When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations

    cs.CL 2026-06 unverdicted novelty 4.0

    Evaluation shows LLMs for healthcare are sensitive to prompt changes, leading to inconsistent and potentially harmful clinical outputs on MedMCQA.