A retrieval-based interleaved visual chain-of-thought method, RIV-CoT, improves VLM answer accuracy by 3.1 points and reasoning accuracy by 4.6 points on a new driving theory VQA benchmark.
Gpt-4 technical report, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios
A retrieval-based interleaved visual chain-of-thought method, RIV-CoT, improves VLM answer accuracy by 3.1 points and reasoning accuracy by 4.6 points on a new driving theory VQA benchmark.