REVIEW 3 cited by
You Only Need Adversarial Supervision for Semantic Image Synthesis
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Despite their recent successes, GAN models for semantic image synthesis still suffer from poor image quality when trained with only adversarial supervision. Historically, additionally employing the VGG-based perceptual loss has helped to overcome this issue, significantly improving the synthesis quality, but at the same time limiting the progress of GAN models for semantic image synthesis. In this work, we propose a novel, simplified GAN model, which needs only adversarial supervision to achieve high quality results. We re-design the discriminator as a semantic segmentation network, directly using the given semantic label maps as the ground truth for training. By providing stronger supervision to the discriminator as well as to the generator through spatially- and semantically-aware discriminator feedback, we are able to synthesize images of higher fidelity with better alignment to their input label maps, making the use of the perceptual loss superfluous. Moreover, we enable high-quality multi-modal image synthesis through global and local sampling of a 3D noise tensor injected into the generator, which allows complete or partial image change. We show that images synthesized by our model are more diverse and follow the color and texture distributions of real images more closely. We achieve an average improvement of $6$ FID and $5$ mIoU points over the state of the art across different datasets using only adversarial supervision.
Forward citations
Cited by 3 Pith papers
-
InstancePin: Instance-Addressable Layout-to-Image Diffusion via Coordinate Pinning
InstancePin adds per-instance coordinate tokens and mask-guided fusion to a frozen layout-to-image diffusion model, improving FID and mIoU on Cityscapes while reducing visual blending of nearby same-category objects.
-
LaMI-GO: Latent Mixture Integration for Goal-Oriented Communications Achieving High Spectrum Efficiency
LaMI-GO sends partially masked visual code indices plus a text caption, and a pre-trained latent diffusion model fills in the masked indices to reconstruct the image at the receiver.
-
Reconciling Semantic Controllability and Diversity for Remote Sensing Image Synthesis with Hybrid Semantic Embedding
HySEGGAN uses mask-derived geometric descriptors plus a segmentation feedback network to improve remote sensing image synthesis and data augmentation.
Discussion (0). Continue with ORCID to comment.