DV-SFT introduces direct vision supervision for visual tokens in MLLMs by auto-labeling from OCR correspondences, outperforming standard SFT on three in-domain and four out-of-domain benchmarks.
arXiv preprint arXiv:2506.09040 , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CV 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
HYDRA-X presents the first unified multimodal model using a single ViT for holistic image-video tokenization, with ablations on attention and compression plus a latent-level editing improvement.
citing papers explorer
-
DV-SFT: Direct Vision Supervision for Fine-Grained Visual Understanding
DV-SFT introduces direct vision supervision for visual tokens in MLLMs by auto-labeling from OCR correspondences, outperforming standard SFT on three in-domain and four out-of-domain benchmarks.
-
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers
HYDRA-X presents the first unified multimodal model using a single ViT for holistic image-video tokenization, with ablations on attention and compression plus a latent-level editing improvement.