A 1.5B unified autoregressive model with separate encoders for generation and understanding reports strong text-to-image and editing scores while running on commodity hardware.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
other 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
other 1polarities
unclear 1representative citing papers
citing papers explorer
-
Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
A 1.5B unified autoregressive model with separate encoders for generation and understanding reports strong text-to-image and editing scores while running on commodity hardware.