Pith. sign in

REVIEW 5 cited by

CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.14217 v2 pith:EMMXUCQB submitted 2022-04-28 cs.CV cs.LG

classification cs.CVcs.LG
keywords generationtext-to-imagecogview2hierarchicalimagestransformersauto-regressiveb-parameter
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The development of the transformer-based text-to-image models are impeded by its slow generation and complexity for high-resolution images. In this work, we put forward a solution based on hierarchical transformers and local parallel auto-regressive generation. We pretrain a 6B-parameter transformer with a simple and flexible self-supervised task, Cross-modal general language model (CogLM), and finetune it for fast super-resolution. The new text-to-image system, CogView2, shows very competitive generation compared to concurrent state-of-the-art DALL-E-2, and naturally supports interactive text-guided editing on images.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A probabilistic overlap measure for object positions yields a human-aligned spatial relationship metric and a training-free generation guidance method for text-to-image models.

  2. Optimizing Multi-Round Enhanced Training in Diffusion Models for Improved Preference Understanding

    cs.CV 2025-04 reject novelty 5.0 of 10

    The VCA framework uses diversity, consistency, and preference rewards to fine-tune a diffusion model with LoRA over multi-round dialogues, reporting improved intent alignment.

  3. Mogo: RQ Hierarchical Causal Transformer for High-Quality 3D Human Motion Generation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Mogo generates 3D human motion from text with a single hierarchical causal transformer and residual vector quantization, reporting a HumanML3D FID of 0.079, the best among GPT-type models.

  4. Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models

    cs.GR 2025-05 conditional novelty 4.0 of 10

    Adding structured metadata such as fabric, sleeve length, and neckline to prompts improves composite quality and ground-truth similarity for most text-to-image models, while slightly reducing prompt-image alignment.

  5. A Multi-Agent Framework for Automated Qinqiang Opera Script Generation Using Large Language Models

    cs.AI 2025-04 reject novelty 4.0 of 10

    A prompt-based multi-agent pipeline generates Qinqiang opera scripts, visuals, and audio, but the reported 0.3-point expert-rated improvement over a vague baseline is not statistically supported.

Pith tools