IoInst evaluates LLMs by asking them to identify which of four instructions generated a given response, and shows that current models often fail, particularly with semantically similar distractors.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Find the Intention of Instruction: Comprehensive Evaluation of Instruction Understanding for Large Language Models
IoInst evaluates LLMs by asking them to identify which of four instructions generated a given response, and shows that current models often fail, particularly with semantically similar distractors.