Pith. sign in

Qwen-Image-2.0 Technical Report

6 Pith papers cite this work. Polarity classification is still indexing.

6 Pith papers citing it
abstract

We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a single framework. Despite recent progress, existing models still struggle with ultra-long text rendering, multilingual typography, high-resolution photorealism, robust instruction following, and efficient deployment, especially in text-rich and compositionally complex scenarios. Qwen-Image-2.0 addresses these challenges by coupling Qwen3-VL as the condition encoder with a Multimodal Diffusion Transformer for joint condition-target modeling, supported by large-scale data curation and a customized multi-stage training pipeline. This enables strong multimodal understanding while preserving flexible generation and editing capabilities. The model supports instructions of up to 1K tokens for generating text-rich content such as slides, posters, infographics, and comics, while significantly improving multilingual text fidelity and typography. It also enhances photorealistic generation with richer details, more realistic textures, and coherent lighting, and follows complex prompts more reliably across diverse styles. Extensive human evaluations show that Qwen-Image-2.0 substantially outperforms previous Qwen-Image models in both generation and editing, marking a step toward more general, reliable, and practical image generation foundation models.

fields

cs.CV 6

years

2026 6

representative citing papers

Qwen-Image-VAE-2.0 Technical Report

cs.CV · 2026-05-13 · unverdicted · novelty 6.0

Qwen-Image-VAE-2.0 achieves state-of-the-art high-compression image reconstruction and superior diffusability for diffusion models, with a new text-rich document benchmark.

Qwen-Image-Flash: Beyond Objective Design

cs.CV · 2026-06-02 · unverdicted · novelty 4.0

Empirical analysis of data, guidance, and task mixture in few-step distillation of Qwen-Image-2.0 produces the Qwen-Image-Flash model with improved performance in unified generation and editing tasks.

Qwen-Image-2.0-RL Technical Report

cs.CV · 2026-06-25 · unverdicted · novelty 2.0

Applies RLHF with composite VLM-based reward models and on-policy distillation to a diffusion model, reporting benchmark gains of +2.61 on Qwen-Image-Bench and Elo improvements of +78/+93.

citing papers explorer

Showing 6 of 6 citing papers.

  • H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks cs.CV · 2026-06-24 · conditional · none · ref 48 · internal anchor

    A diffusion-based hairstyle transfer method that uses a region-specific training loss to make cross-attention produce a source-aligned hair mask for pose-robust inpainting.

  • WeGenBench: A Multidimensional Diagnostic Benchmark towards Text-to-Image Model Optimization cs.CV · 2026-06-18 · unverdicted · none · ref 2 · internal anchor

    WeGenBench provides 4000 bilingual prompts with scene and tag annotations plus VLM-derived metrics to locate specific deficiencies in text-to-image models.

  • Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation cs.CV · 2026-05-27 · unverdicted · none · ref 17 · internal anchor

    Qwen-Image-Bench introduces a hierarchical creator-centric benchmark with 1000 prompts, 23 sub-capabilities, and a Q-Judger model that scores images on 56 verifiable facets to distinguish T2I models on fidelity and creativity.

  • Qwen-Image-VAE-2.0 Technical Report cs.CV · 2026-05-13 · unverdicted · none · ref 19 · internal anchor

    Qwen-Image-VAE-2.0 achieves state-of-the-art high-compression image reconstruction and superior diffusability for diffusion models, with a new text-rich document benchmark.

  • Qwen-Image-Flash: Beyond Objective Design cs.CV · 2026-06-02 · unverdicted · none · ref 17 · internal anchor

    Empirical analysis of data, guidance, and task mixture in few-step distillation of Qwen-Image-2.0 produces the Qwen-Image-Flash model with improved performance in unified generation and editing tasks.

  • Qwen-Image-2.0-RL Technical Report cs.CV · 2026-06-25 · unverdicted · none · ref 18 · internal anchor

    Applies RLHF with composite VLM-based reward models and on-policy distillation to a diffusion model, reporting benchmark gains of +2.61 on Qwen-Image-Bench and Elo improvements of +78/+93.