PromptSET is a benchmark of 11,469 questions with nine LLM-generated rephrasings each, and current classifiers and self-evaluation methods predict prompt answerability poorly, especially on multi-hop questions.
Title resolution pending
1 Pith paper cite this work, alongside 20 external citations. Polarity classification is still indexing.
1
Pith paper citing it
20
external citations · OpenAlex
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Benchmarking Prompt Sensitivity in Large Language Models
PromptSET is a benchmark of 11,469 questions with nine LLM-generated rephrasings each, and current classifiers and self-evaluation methods predict prompt answerability poorly, especially on multi-hop questions.