K-FinHallu is the first multi-turn Korean financial RAG hallucination benchmark; frontier LLMs struggle especially on justified abstention while an 8B fine-tuned model reaches competitive performance.
Lewis and 1 others
3 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
MTR-Suite offers an LLM-based auditor, a low-cost multi-agent synthesis pipeline using greedy traversal clustering, and a new general-domain benchmark with superior discriminative power for conversational retrieval.
RAG-DIVE uses an LLM to dynamically generate, validate, and evaluate multi-turn dialogues for assessing RAG system performance in interactive settings.
citing papers explorer
-
K-FinHallu: A Hallucination Detection Benchmark for Multi-Turn RAG in Korean Finance
K-FinHallu is the first multi-turn Korean financial RAG hallucination benchmark; frontier LLMs struggle especially on justified abstention while an 8B fine-tuned model reaches competitive performance.
-
MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks
MTR-Suite offers an LLM-based auditor, a low-cost multi-agent synthesis pipeline using greedy traversal clustering, and a new general-domain benchmark with superior discriminative power for conversational retrieval.
-
RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation
RAG-DIVE uses an LLM to dynamically generate, validate, and evaluate multi-turn dialogues for assessing RAG system performance in interactive settings.