REVIEW 13 cited by
What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) changed the way we design and interact with software systems. Their ability to process and extract information from text has drastically improved productivity in a number of routine tasks. Developers that want to include these models in their software stack, however, face a dreadful challenge: debugging LLMs' inconsistent behavior across minor variations of the prompt. We therefore introduce two metrics for classification tasks, namely sensitivity and consistency, which are complementary to task performance. First, sensitivity measures changes of predictions across rephrasings of the prompt, and does not require access to ground truth labels. Instead, consistency measures how predictions vary across rephrasings for elements of the same class. We perform an empirical comparison of these metrics on text classification tasks, using them as guideline for understanding failure modes of the LLM. Our hope is that sensitivity and consistency will be helpful to guide prompt engineering and obtain LLMs that balance robustness with performance.
Forward citations
Cited by 13 Pith papers
-
From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models
A weak-supervision pipeline using LLM-extracted concepts and rank fusion creates silver-standard patent-to-SDG labels that recover known citation-derived associations and show high network modularity.
-
Reinforcement Learning for Machine Learning Engineering Agents
RL-trained Qwen2.5-3B outperforms prompted Claude-3.5-Sonnet and GPT-4o on 12 MLEBench tasks by an average of 22% and 24%, using two targeted RL modifications.
-
Adaptive Prompting: Ad-hoc Prompt Composition for Social Bias Detection
On StereoSet and SBIC, a trained DeBERTa encoder that selects per-input prompt compositions from 64 options raises macro F1 above every fixed composition, but on CobraFrames it falls below the best fixed composition.
-
Linguistic Features Extracted by GPT-4 Improve Alzheimer's Disease Detection based on Spontaneous Speech
GPT-4 ratings of five dementia-related language symptoms, added to 40 standard linguistic features, improve automatic Alzheimer's detection from spontaneous speech transcripts, reaching AUROC 0.931 on ADReSS.
-
When the Code Autopilot Breaks: Why LLMs Falter in Embedded Machine Learning
LLM-based sketch generation for embedded ML is fragile, with success rates below 40%, and prompt structure alone can swing outcomes from 15% to 30%.
-
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
Across 156 configurations on Persian medical board questions, Chain-of-Thought prompting raised accuracy while increasing overconfidence, and emotional prompting inflated confidence without accuracy gains.
-
Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets
AI agents in future labor markets will need metacognitive and strategic reasoning because incomplete information creates adverse selection, moral hazard, and reputation effects.
-
Position: Contextual Integrity is Inadequately Applied to Language Models
Many prior studies that apply Contextual Integrity to language models omit the theory's core tenets, which can make their privacy conclusions unreliable.
-
AI-Assisted Data Extraction for Systematic Reviews in Education
Free LLM APIs extract systematic review data with only about 62 to 72 percent exact agreement with human coding, so a human-in-the-loop tool (AIDE) is proposed to validate every extracted item.
-
Leveraging Prior Experience: An Expandable Auxiliary Knowledge Base for Text-to-SQL
LPE-SQL improves text-to-SQL accuracy on BIRD by retrieving from dynamically grown correct and mistake notebooks, but its main gains come from feeding ground-truth answers into those notebooks during evaluation.
-
Position: Intelligent Coding Systems Should Write Programs with Justifications
A position paper advocating that intelligent coding systems should accompany code with justified explanations that are cognitively aligned and semantically faithful.
-
A Conceptual Framework for Requirements Engineering of Pretrained-Model-Enabled Systems
A conceptual framework reorganizes requirements engineering for pretrained-model-enabled systems into six activities, based on identified challenges of opaque capabilities, context sensitivity, and continuous evolution.
-
CEA-LIST at CheckThat! 2025: Evaluating LLMs as Detectors of Bias and Opinion in Text
Few-shot prompted LLMs rivaled fine-tuned smaller models in multilingual subjectivity detection, winning the Arabic and Polish tracks of CheckThat! 2025.
Discussion (0). Continue with ORCID to comment.