MTLA is a training-free, post-hoc confidence score for multimodal LLM localization that restricts attention aggregation to the model's own predicted region and tokens, substantially improving hallucination detection and re-ranking across image, video, and audio.
Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention
MTLA is a training-free, post-hoc confidence score for multimodal LLM localization that restricts attention aggregation to the model's own predicted region and tokens, substantially improving hallucination detection and re-ranking across image, video, and audio.