Pith. sign in

REVIEW 13 cited by

What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.12334 v4 pith:OSYBTQOA submitted 2024-06-18 cs.LG cs.SE

classification cs.LGcs.SE
keywords consistencyllmspromptsensitivityacrosstasksclassificationengineering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) changed the way we design and interact with software systems. Their ability to process and extract information from text has drastically improved productivity in a number of routine tasks. Developers that want to include these models in their software stack, however, face a dreadful challenge: debugging LLMs' inconsistent behavior across minor variations of the prompt. We therefore introduce two metrics for classification tasks, namely sensitivity and consistency, which are complementary to task performance. First, sensitivity measures changes of predictions across rephrasings of the prompt, and does not require access to ground truth labels. Instead, consistency measures how predictions vary across rephrasings for elements of the same class. We perform an empirical comparison of these metrics on text classification tasks, using them as guideline for understanding failure modes of the LLM. Our hope is that sensitivity and consistency will be helpful to guide prompt engineering and obtain LLMs that balance robustness with performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A weak-supervision pipeline using LLM-extracted concepts and rank fusion creates silver-standard patent-to-SDG labels that recover known citation-derived associations and show high network modularity.

  2. Reinforcement Learning for Machine Learning Engineering Agents

    cs.LG 2025-09 conditional novelty 6.0 of 10

    RL-trained Qwen2.5-3B outperforms prompted Claude-3.5-Sonnet and GPT-4o on 12 MLEBench tasks by an average of 22% and 24%, using two targeted RL modifications.

  3. Adaptive Prompting: Ad-hoc Prompt Composition for Social Bias Detection

    cs.CL 2025-02 conditional novelty 6.0 of 10

    On StereoSet and SBIC, a trained DeBERTa encoder that selects per-input prompt compositions from 64 options raises macro F1 above every fixed composition, but on CobraFrames it falls below the best fixed composition.

  4. Linguistic Features Extracted by GPT-4 Improve Alzheimer's Disease Detection based on Spontaneous Speech

    cs.CL 2024-12 conditional novelty 6.0 of 10

    GPT-4 ratings of five dementia-related language symptoms, added to 40 standard linguistic features, improve automatic Alzheimer's detection from spontaneous speech transcripts, reaching AUROC 0.931 on ADReSS.

  5. When the Code Autopilot Breaks: Why LLMs Falter in Embedded Machine Learning

    cs.SE 2025-09 conditional novelty 5.0 of 10

    LLM-based sketch generation for embedded ML is fragile, with success rates below 40%, and prompt structure alone can swing outcomes from 15% to 30%.

  6. Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs

    cs.CY 2025-05 conditional novelty 5.0 of 10

    Across 156 configurations on Persian medical board questions, Chain-of-Thought prompting raised accuracy while increasing overconfidence, and emotional prompting inflated confidence without accuracy gains.

  7. Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets

    cs.AI 2025-05 conditional novelty 5.0 of 10

    AI agents in future labor markets will need metacognitive and strategic reasoning because incomplete information creates adverse selection, moral hazard, and reputation effects.

  8. Position: Contextual Integrity is Inadequately Applied to Language Models

    cs.CY 2025-01 conditional novelty 5.0 of 10

    Many prior studies that apply Contextual Integrity to language models omit the theory's core tenets, which can make their privacy conclusions unreliable.

  9. AI-Assisted Data Extraction for Systematic Reviews in Education

    cs.HC 2025-01 conditional novelty 5.0 of 10

    Free LLM APIs extract systematic review data with only about 62 to 72 percent exact agreement with human coding, so a human-in-the-loop tool (AIDE) is proposed to validate every extracted item.

  10. Leveraging Prior Experience: An Expandable Auxiliary Knowledge Base for Text-to-SQL

    cs.CL 2024-11 reject novelty 5.0 of 10

    LPE-SQL improves text-to-SQL accuracy on BIRD by retrieving from dynamically grown correct and mistake notebooks, but its main gains come from feeding ground-truth answers into those notebooks during evaluation.

  11. Position: Intelligent Coding Systems Should Write Programs with Justifications

    cs.SE 2025-08 conditional novelty 4.0 of 10

    A position paper advocating that intelligent coding systems should accompany code with justified explanations that are cognitively aligned and semantically faithful.

  12. A Conceptual Framework for Requirements Engineering of Pretrained-Model-Enabled Systems

    cs.SE 2025-07 conditional novelty 4.0 of 10

    A conceptual framework reorganizes requirements engineering for pretrained-model-enabled systems into six activities, based on identified challenges of opaque capabilities, context sensitivity, and continuous evolution.

  13. CEA-LIST at CheckThat! 2025: Evaluating LLMs as Detectors of Bias and Opinion in Text

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Few-shot prompted LLMs rivaled fine-tuned smaller models in multilingual subjectivity detection, winning the Arabic and Polish tracks of CheckThat! 2025.

Pith tools