Pith. sign in

REVIEW 2 cited by

N\"UWA: Visual Synthesis Pre-training for Neural visUal World creAtion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.12417 v1 pith:2DJQ5WXZ submitted 2021-11-24 cs.CV cs.AI

classification cs.CVcs.AI
keywords visualdatatasksvideogenerationimageimagessynthesis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents a unified multimodal pre-trained model called N\"UWA that can generate new or manipulate existing visual data (i.e., images and videos) for various visual synthesis tasks. To cover language, image, and video at the same time for different scenarios, a 3D transformer encoder-decoder framework is designed, which can not only deal with videos as 3D data but also adapt to texts and images as 1D and 2D data, respectively. A 3D Nearby Attention (3DNA) mechanism is also proposed to consider the nature of the visual data and reduce the computational complexity. We evaluate N\"UWA on 8 downstream tasks. Compared to several strong baselines, N\"UWA achieves state-of-the-art results on text-to-image generation, text-to-video generation, video prediction, etc. Furthermore, it also shows surprisingly good zero-shot capabilities on text-guided image and video manipulation tasks. Project repo is https://github.com/microsoft/NUWA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation

    cs.CV 2024-12 conditional novelty 3.0 of 10

    A survey of long video generation that groups methods into frame-by-frame, planned-segment, and all-at-once approaches, but is weakened by inconsistent data tables.

  2. Movie Gen: SWOT Analysis of Meta's Generative AI Foundation Model for Transforming Media Generation, Advertising, and Entertainment Industries

    cs.AI 2024-12 unverdicted novelty 2.0 of 10

    A SWOT analysis of Meta's Movie Gen that restates vendor-reported features without new evidence.

Pith tools