Pith. sign in

REVIEW 3 cited by

Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.03509 v3 pith:ZSQQYUQI submitted 2023-05-04 cs.CL cs.AIcs.HCcs.LG

classification cs.CLcs.AIcs.HCcs.LG
keywords diffusionexplainerstablecomplexgenerationimageimagesnon-experts
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Diffusion-based generative models' impressive ability to create convincing images has garnered global attention. However, their complex structures and operations often pose challenges for non-experts to grasp. We present Diffusion Explainer, the first interactive visualization tool that explains how Stable Diffusion transforms text prompts into images. Diffusion Explainer tightly integrates a visual overview of Stable Diffusion's complex structure with explanations of the underlying operations. By comparing image generation of prompt variants, users can discover the impact of keyword changes on image generation. A 56-participant user study demonstrates that Diffusion Explainer offers substantial learning benefits to non-experts. Our tool has been used by over 10,300 users from 124 countries at https://poloclub.github.io/diffusion-explainer/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ChannelExplorer: Exploring Class Separability Through Activation Channel Visualization

    cs.GR 2025-05 conditional novelty 6.0 of 10

    ChannelExplorer turns activation channel summaries into scatterplots, Jaccard similarity matrices, and heatmaps, giving users a way to explore class separability in neural networks.

  2. Mapping the Mind of an Instruction-based Image Editing using SMILE

    cs.AI 2024-12 reject novelty 4.0 of 10

    SMILE applies LIME-style prompt perturbation with image-embedding distances to create word-level heatmaps for instruction-based image editing models.

  3. Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A survey maps the field of MLLM explainability and interpretability into data, model, and training and inference perspectives.

Pith tools