Pith. sign in

REVIEW 2 cited by

Schr\"{o}dinger's Bat: Diffusion Models Sometimes Generate Polysemous Words in Superposition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.13095 v1 pith:5GD32Z6O submitted 2022-11-23 cs.CL

classification cs.CL
keywords diffusionimagesmeaningsmodelspolysemoussuperpositionwordwords
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work has shown that despite their impressive capabilities, text-to-image diffusion models such as DALL-E 2 (Ramesh et al., 2022) can display strange behaviours when a prompt contains a word with multiple possible meanings, often generating images containing both senses of the word (Rassin et al., 2022). In this work we seek to put forward a possible explanation of this phenomenon. Using the similar Stable Diffusion model (Rombach et al., 2022), we first show that when given an input that is the sum of encodings of two distinct words, the model can produce an image containing both concepts represented in the sum. We then demonstrate that the CLIP encoder used to encode prompts (Radford et al., 2021) encodes polysemous words as a superposition of meanings, and that using linear algebraic techniques we can edit these representations to influence the senses represented in the generated images. Combining these two findings, we suggest that the homonym duplication phenomenon described by Rassin et al. (2022) is caused by diffusion models producing images representing both of the meanings that are present in superposition in the encoding of a polysemous word.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Where did the ambiguity go? Examining how multimodal models interpret polysemous words

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Across 17 image and 15 text models, generated images settle on far fewer senses of ambiguous words than generated sentences, and both fall well short of human diversity.

  2. Enhancing Generalization in Data-free Quantization via Mixup-class Prompting

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Using two class labels in text prompts to generate synthetic calibration images improves data-free post-training quantization accuracy, especially in low-bit settings.

Pith tools