Benchmark construction artifacts in hallucination detection corpora allow naive text-similarity baselines to achieve near-perfect scores, and controlled evaluations show most methods perform near chance except SAPLMA and the new DRIFT probe.
MedHallBench: a benchmark for hallucination detection in med- ical VLMs with reinforcement learning-assisted annotation,
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2verdicts
UNVERDICTED 2representative citing papers
A literature synthesis that unifies hallucination taxonomies across medical imaging modalities, finds general-purpose foundation models hallucinate less than specialized ones, and maps mitigation to FDA lifecycle frameworks.
citing papers explorer
-
PARALLAX: Separating Genuine Hallucination Detection from Benchmark Construction Artifacts
Benchmark construction artifacts in hallucination detection corpora allow naive text-similarity baselines to achieve near-perfect scores, and controlled evaluations show most methods perform near chance except SAPLMA and the new DRIFT probe.
-
Hallucination in Medical Imaging AI: A Cross-Modality Analytical Framework for Taxonomy, Detection, and Mitigation under Regulatory Constraints
A literature synthesis that unifies hallucination taxonomies across medical imaging modalities, finds general-purpose foundation models hallucinate less than specialized ones, and maps mitigation to FDA lifecycle frameworks.