Pith. sign in

REVIEW 5 cited by

Visual Anagrams: Generating Multi-View Optical Illusions with Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.17919 v2 pith:V4VMONQU submitted 2023-11-29 cs.CV

classification cs.CV
keywords illusionsdiffusionimagemethodviewsvisualanagramsappearance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We address the problem of synthesizing multi-view optical illusions: images that change appearance upon a transformation, such as a flip or rotation. We propose a simple, zero-shot method for obtaining these illusions from off-the-shelf text-to-image diffusion models. During the reverse diffusion process, we estimate the noise from different views of a noisy image, and then combine these noise estimates together and denoise the image. A theoretical analysis suggests that this method works precisely for views that can be written as orthogonal transformations, of which permutations are a subset. This leads to the idea of a visual anagram--an image that changes appearance under some rearrangement of pixels. This includes rotations and flips, but also more exotic pixel permutations such as a jigsaw rearrangement. Our approach also naturally extends to illusions with more than two views. We provide both qualitative and quantitative results demonstrating the effectiveness and flexibility of our method. Please see our project webpage for additional visualizations and results: https://dangeng.github.io/visual_anagrams/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lossy Compression with Pretrained Diffusion Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A complete, zero-shot implementation of the DiffC algorithm lets pretrained Stable Diffusion models act as lossy image compressors at ultra-low bitrates.

  2. Illusion3D: 3D Multiview Illusion with 2D Diffusion Priors

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Illusion3D generates 3D objects with multicolor textures that reveal different pictures from different viewpoints, using a 2D text-to-image diffusion model and score-distillation optimization.

  3. Diffusion-based Visual Anagram as Multi-task Learning

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A diffusion-based method generates visual anagrams by treating each viewpoint as a task and adding anti-segregation, noise-balancing, and variance-rectification steps.

  4. Making Images from Images: Interleaving Denoising and Transformation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A diffusion-based system that learns, during generation, the tile rearrangement that turns a fixed source image into a new image described by a text prompt.

  5. Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A new benchmark plus a blur-based filter that make vision-language models better at recognizing hidden classes in synthetic pareidolia images.

Pith tools