MARCH benchmark shows state-of-the-art models struggle on ambiguous multi-hop QA, and the proposed CLARION framework improves performance by decoupling ambiguity planning from evidence-driven reasoning.
History”,“Geography&Places
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CL 2verdicts
UNVERDICTED 2representative citing papers
Each tested LLM shows its own characteristic unreliability when engaging in repair during extended math-question dialogues.
citing papers explorer
-
MARCH: Evaluating the Intersection of Ambiguity Interpretation and Multi-hop Inference
MARCH benchmark shows state-of-the-art models struggle on ambiguous multi-hop QA, and the proposed CLARION framework improves performance by decoupling ambiguity planning from evidence-driven reasoning.
-
Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs
Each tested LLM shows its own characteristic unreliability when engaging in repair during extended math-question dialogues.