POINTS-Reader uses a two-stage synthetic-data warm-up plus iterative self-improvement with rule-based filtering to train a 3B vision-language model for document conversion, outperforming larger models on OmniDocBench and Fox.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
POINTS-Reader uses a two-stage synthetic-data warm-up plus iterative self-improvement with rule-based filtering to train a 3B vision-language model for document conversion, outperforming larger models on OmniDocBench and Fox.