An ensemble of five LLM evaluators plus conformal prediction labeled surgical instructions as ambiguous or clear with 70% (Llama 3.2 11B) and 82.5% (Gemma 3 12B) accuracy, measured in-sample on the 40-instruction calibration set.
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We introduce AmbigNLG, a novel task designed to tackle the challenge of task ambiguity in instructions for Natural Language Generation (NLG). Ambiguous instructions often impede the performance of Large Language Models (LLMs), especially in complex NLG tasks. To tackle this issue, we propose an ambiguity taxonomy that categorizes different types of instruction ambiguities and refines initial instructions with clearer specifications. Accompanying this task, we present AmbigSNI-NLG, a dataset comprising 2,500 instances annotated to facilitate research in AmbigNLG. Through comprehensive experiments with state-of-the-art LLMs, we demonstrate that our method significantly enhances the alignment of generated text with user expectations, achieving up to a 15.02-point increase in ROUGE scores. Our findings highlight the critical importance of addressing task ambiguity to fully harness the capabilities of LLMs in NLG tasks. Furthermore, we confirm the effectiveness of our method in practical settings involving interactive ambiguity mitigation with users, underscoring the benefits of leveraging LLMs for interactive clarification.
citation-role summary
citation-polarity summary
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
LLM-based ambiguity detection in natural language instructions for collaborative surgical robots
An ensemble of five LLM evaluators plus conformal prediction labeled surgical instructions as ambiguous or clear with 70% (Llama 3.2 11B) and 82.5% (Gemma 3 12B) accuracy, measured in-sample on the 40-instruction calibration set.