ScreenAnnotator introduces a unified annotation atom schema, on-policy Bayesian verification loop, and template-driven synthesis to generate reusable multi-task reasoning data for VLMs, reporting nearly 100% accept rate on flowcharts, 77% on GUI screenshots, and a 35.1-point accuracy gain after fine
arXiv preprint arXiv:2311.17076 , year=
4 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 4representative citing papers
SeProD is a plug-and-play self-prophetic decoding framework that combines pre- and post-training LVLM capabilities via probability-based sampling to improve coherent visual search and multi-step reasoning.
Representations learned by large AI models are converging toward a shared statistical model of reality.
UnAC improves LMM performance on visual reasoning benchmarks by combining adaptive visual prompting, image abstraction, and gradual self-checking.
citing papers explorer
-
From Bounding Boxes to Visual Reasoning: An On-Policy Data Annotation Tool for Vision-Language Models
ScreenAnnotator introduces a unified annotation atom schema, on-policy Bayesian verification loop, and template-driven synthesis to generate reusable multi-task reasoning data for VLMs, reporting nearly 100% accept rate on flowcharts, 77% on GUI screenshots, and a 35.1-point accuracy gain after fine
-
Self-Prophetic Decoding to Unlock Visual Search in LVLMs
SeProD is a plug-and-play self-prophetic decoding framework that combines pre- and post-training LVLM capabilities via probability-based sampling to improve coherent visual search and multi-step reasoning.
-
The Platonic Representation Hypothesis
Representations learned by large AI models are converging toward a shared statistical model of reality.
-
UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning
UnAC improves LMM performance on visual reasoning benchmarks by combining adaptive visual prompting, image abstraction, and gradual self-checking.