CaVe-VLM-CoT is a closed-loop agentic-RAG framework with Extractor, Retriever, Solver, Citation Injector and Verifier stages plus 23 metrics anchored by CaVeScore that reports 87.1% accuracy on ScienceQA and 55.2% on MMMU without model changes.
arXiv preprint arXiv:2508.00378 , year=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework
CaVe-VLM-CoT is a closed-loop agentic-RAG framework with Extractor, Retriever, Solver, Citation Injector and Verifier stages plus 23 metrics anchored by CaVeScore that reports 87.1% accuracy on ScienceQA and 55.2% on MMMU without model changes.