Pith. sign in

REVIEW 1 cited by

GANs N' Roses: Stable, Controllable, Diverse Image to Image Translation (works for videos too!)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.06561 v1 pith:YWYW7HGW submitted 2021-06-11 cs.CV cs.GRcs.LG

classification cs.CVcs.GRcs.LG
keywords animecontentdiverseimagecodestyleadversarialextensive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We show how to learn a map that takes a content code, derived from a face image, and a randomly chosen style code to an anime image. We derive an adversarial loss from our simple and effective definitions of style and content. This adversarial loss guarantees the map is diverse -- a very wide range of anime can be produced from a single content code. Under plausible assumptions, the map is not just diverse, but also correctly represents the probability of an anime, conditioned on an input face. In contrast, current multimodal generation procedures cannot capture the complex styles that appear in anime. Extensive quantitative experiments support the idea the map is correct. Extensive qualitative results show that the method can generate a much more diverse range of styles than SOTA comparisons. Finally, we show that our formalization of content and style allows us to perform video to video translation without ever training on videos.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Advancing Facial Stylization through Semantic Preservation Constraint and Pseudo-Paired Supervision

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A StyleGAN fine-tuning recipe with semantic preservation and multi-level pseudo-paired supervision yields higher-fidelity facial stylization, plus free multimodal and reference-guided variants.

Pith tools