Fine-tuning Qwen2.5-7B on GPT-4o-generated perturbed multi-hop QA data yields recall 0.938 vs GPT-4o's 0.710 on RAGTruth hallucination detection, with lower precision (0.366 vs 0.446).
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Osiris: A Lightweight Open-Source Hallucination Detection System
Fine-tuning Qwen2.5-7B on GPT-4o-generated perturbed multi-hop QA data yields recall 0.938 vs GPT-4o's 0.710 on RAGTruth hallucination detection, with lower precision (0.366 vs 0.446).