Text-to-image models systematically fail to apply an object's intrinsic frame of reference: mean final accuracy drops from 26.5% on camera-view prompts to 15.4% on matched frame-of-reference prompts, and the failure peaks when the anchor's left or right is reversed relative to the image.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Can Text-to-Image Models Draw from the Right Frame of Reference?
Text-to-image models systematically fail to apply an object's intrinsic frame of reference: mean final accuracy drops from 26.5% on camera-view prompts to 15.4% on matched frame-of-reference prompts, and the failure peaks when the anchor's left or right is reversed relative to the image.