S2CAN generates self-contained direct and indirect memories (hints and question-hint pairs) from surgical images, then reasons over them to answer surgical VQA questions, reporting SOTA results on EndoVis-18, EndoVis-17, and Cholec80 benchmarks.
Lawrence Zitnick, and Devi Parikh
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry
S2CAN generates self-contained direct and indirect memories (hints and question-hint pairs) from surgical images, then reasons over them to answer surgical VQA questions, reporting SOTA results on EndoVis-18, EndoVis-17, and Cholec80 benchmarks.