A frozen DiT plus task LoRA and a ~33K-parameter token-local linear head reads out pixel-space dense fields and sets SOTA on matting, KITTI depth, and referring segmentation while running up to 2.48 imes faster than edit-plus-decode baselines.
In: Pro- ceedings of the IEEE/CVF International Confer- ence on Computer Vision
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models
A frozen DiT plus task LoRA and a ~33K-parameter token-local linear head reads out pixel-space dense fields and sets SOTA on matting, KITTI depth, and referring segmentation while running up to 2.48 imes faster than edit-plus-decode baselines.