Pith. sign in

REVIEW 2 cited by

A Cat Is A Cat (Not A Dog!): Unraveling Information Mix-ups in Text-to-Image Encoders through Causal Analysis and Embedding Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.00321 v5 pith:SUZUOXKK submitted 2024-10-01 cs.CV

classification cs.CV
keywords embeddinginformationtextaddressinganalysisbalancecausalcontributes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper analyzes the impact of causal manner in the text encoder of text-to-image (T2I) diffusion models, which can lead to information bias and loss. Previous works have focused on addressing the issues through the denoising process. However, there is no research discussing how text embedding contributes to T2I models, especially when generating more than one object. In this paper, we share a comprehensive analysis of text embedding: i) how text embedding contributes to the generated images and ii) why information gets lost and biases towards the first-mentioned object. Accordingly, we propose a simple but effective text embedding balance optimization method, which is training-free, with an improvement of 125.42% on information balance in stable diffusion. Furthermore, we propose a new automatic evaluation metric that quantifies information loss more accurately than existing methods, achieving 81% concordance with human assessments. This metric effectively measures the presence and accuracy of objects, addressing the limitations of current distribution scores like CLIP's text-image similarities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Detail++: Training-Free Detail Enhancer for T2I Diffusion Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Detail++ uses progressive multi-branch prompt injection and test-time attention optimization to improve attribute binding in text-to-image generation.

  2. FreeCond: Free Lunch in the Input Conditions of Text-Guided Inpainting

    cs.CV 2024-11 conditional novelty 6.0 of 10

    FreeCond adjusts only the image and mask inputs of Stable Diffusion Inpainting, improving prompt adherence and mask fitting without training or extra compute.

Pith tools