Pith. sign in

Paper Citation Record · LEDGER

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps

As of 22 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2411.15236.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15236 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:08:54.669287Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:45:53.297589Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T15:58:37.263440Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 860654df-32b1-470a-a7be-77ce9b923569 · outbound

This paper cites A-star: Test-time attention segregation and retention for text-to-image synthesis.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps A-star: Test-time attention segregation and retention for text-to-image synthesis

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.587696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T15:08:53.220813Z digest=sha256:2836caf96c1232714f7802b8b9343deb91f0bfd78bb976a09826d0d50a3f2f79

Observation 00464a84-ec3f-481f-9cba-3fb45053ed2d · outbound

This paper cites AlignIT: Enhancing Prompt Alignment in Customization of Text-to-Image Models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps AlignIT: Enhancing Prompt Alignment in Customization of Text-to-Image Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:08:55.655150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T15:08:53.416975Z digest=sha256:b853ab92c9c05119eafd9d4b0d68942fed2e1d676902c0d142e82168d0d89e8c

Observation f3a26ef5-c45c-459a-a889-f967ed3a35fd · outbound

This paper cites Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.510085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T15:08:53.538102Z digest=sha256:2d549898d4ad323390653e77941ae68112304cc726d842d135fb2ce2586e3db9

Observation 86b56c1a-ed34-4e6f-a549-f9508e3e5072 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.586669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.586669Z digest=sha256:89e3b3ecca30c2c2ade8d2e26db39f58dca12b91128d43b2adf6f41b50afc04b

Observation c1746190-6bc2-4565-81e2-59ffb069b285 · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.596085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.596085Z digest=sha256:80e37429518eb5ae87a3231060b6aebc48138fd6bab059ccc0283a815f1673b0

Observation 338f1936-24c5-4ed3-94a4-fa7ee20f7b0b · outbound

This paper cites Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.459247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T15:08:53.629726Z digest=sha256:0019c730d8ee36450a5fff349fe781aba35cdab54ee958926967e35187a23bc9

Observation a614f833-0c21-4618-8e76-d663731cda2e · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.675405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.675405Z digest=sha256:3bcaf1634bad071489d399bca4a22d9ae4b9346435137c346b3a085d9c5b1be9

Observation 38e3d833-4cdc-45bd-8fd6-117feeac26d1 · outbound

This paper cites Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.683502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.683502Z digest=sha256:fce53d0bd19fb5b8e8e653eea4e78ba08f4e86937d5121aeb52fb76008a4a611

Observation 735bb36e-7c43-4f19-bb83-85f56785282f · outbound

This paper cites When Attention Sink Emerges in Language Models: An Empirical View.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps When Attention Sink Emerges in Language Models: An Empirical View

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.730797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.730797Z digest=sha256:1ae6e41ace5a518199911ffe0d6adc2cde977a4e17534a8e56876afd72c098ef

Observation 9f994c93-3122-47c8-bd26-2ca7ec81dc13 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.768421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.768421Z digest=sha256:ebeee766315a46a1a3ed6e7d94399b111bd4123f83b19a0d94d55067fe39f7bf

Observation e0a854e5-19ca-4449-8159-8bf6b8d2fed6 · outbound

This paper cites spacy 2: Natural lan- guage understanding with bloom embeddings, convolutional neural networks and incremental parsing.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps spacy 2: Natural lan- guage understanding with bloom embeddings, convolutional neural networks and incremental parsing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.362672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T15:08:53.783216Z digest=sha256:962371cf4150a8f488446ac6a7bd04ee8f1f6f5855a3832039334499df335654

Observation ce474741-af47-465b-bc87-cbc72ca61bd7 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.821243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.821243Z digest=sha256:455bd36d5685ce38bd60a071e4c8bad7183c8b576f280b8b5e5d383af2f789f7

Observation 7ffb49e3-94ef-49ad-a930-5195b5240431 · outbound

This paper cites Tifa: Accu- rate and interpretable text-to-image faithfulness evaluation with question answering.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Tifa: Accu- rate and interpretable text-to-image faithfulness evaluation with question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.288975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T15:08:53.842397Z digest=sha256:178074a04cce2dbc9cb604a4fbcf74a6d30dc87fbaa0e3948ce1dc5f32dfd618

Observation 4139fcd9-6cf9-4359-9bca-762418452722 · outbound

This paper cites MC$^2$: Multi-concept Guidance for Customized Multi-concept Generation.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps MC$^2$: Multi-concept Guidance for Customized Multi-concept Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.854867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.854867Z digest=sha256:8114710c02d45c83de856a640b6f8755b059869300ad3b6dbaf2a2ad75fc2bb1

Observation 166d9dcb-4ec2-453d-b425-f61bd04fea16 · outbound

This paper cites Dense text-to-image generation with attention modulation.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Dense text-to-image generation with attention modulation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.879934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.879934Z digest=sha256:c9ab2a5f1b542d26125f76e9ef28ee9dc246cfeabb005a6570b2d706616f58a4

Observation 9be8d80a-c345-45a0-a1d6-eb7b3402a058 · outbound

This paper cites mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.954810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.954810Z digest=sha256:606aeccf731ca4f77f3e0cc4532efe0fddcb141d7fd532b7fc4e94cdba6c8321

Observation 978e7c52-b3e2-417e-9378-8a35852e42b8 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.982683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.982683Z digest=sha256:cd30a95b608401b3e9333bf5397e36d2a1106467ac03d2d593765c49cf12b911

Observation cd0d00d4-ea19-4ec4-9c31-f439499692a8 · outbound

This paper cites Microsoft coco: Common objects in context.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Microsoft coco: Common objects in context

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.993512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.993512Z digest=sha256:4660f87a1eae7bb01a0bf2fe2b0b4404a4e2a9bfa02c7bf33ebd6dc67d71bffb

Observation 37b1587b-5b1e-4954-a3ad-5feb3b425a7a · outbound

This paper cites Improving Text-to-Image Consistency via Automatic Prompt Optimization.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Improving Text-to-Image Consistency via Automatic Prompt Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.008711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.008711Z digest=sha256:ba584a0f5f1f3d73bb676af19bab7ff4cfcefbd0669d1859455a564d9f8e966e

Observation c62ebd91-f847-4e46-a36a-22593e22af9e · outbound

This paper cites Attention Overlap Is Responsible for The Entity Missing Problem in Text-to-image Diffusion Models!.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Attention Overlap Is Responsible for The Entity Missing Problem in Text-to-image Diffusion Models!

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:08:55.214378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T15:08:54.067179Z digest=sha256:50176ad6ce8273f82a80945e6b44522fad18675ea45df8105be824a52b4dfde5

Observation 22456177-2924-4e5a-80e3-d21454ed916a · outbound

This paper cites Conform: Contrast is all you need for high- fidelity text-to-image diffusion models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Conform: Contrast is all you need for high- fidelity text-to-image diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.127872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T15:08:54.125524Z digest=sha256:8fb150040beddca4de5c3950bb1aa43c15cb281d6c2ac960c58238bfa53884ad

Observation 70120ce3-9e5c-4932-97f2-7d689dec5a89 · outbound

This paper cites Openai gpt-3 api [gpt-3.5-turbo], 2024.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Openai gpt-3 api [gpt-3.5-turbo], 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.003417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T15:08:54.138455Z digest=sha256:a3e5eab4226b1a633ccff4a4451094662d8f35cd4043e4e4615056a4a61d8659

Observation 1b51419f-9f9c-4f71-bb8d-f1216a926d55 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.150477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.150477Z digest=sha256:f2d4b4b37a558748ce720b371eeae2d4387c8d40e6ebce0133402cd1fe4db76d

Observation f7c116de-be8d-4031-88e7-7e733ee9434a · outbound

This paper cites Not All Noises Are Created Equally:Diffusion Noise Selection and Optimization.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Not All Noises Are Created Equally:Diffusion Noise Selection and Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.173401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.173401Z digest=sha256:ba2a5bc414fa2d1663a15ea7bc1cb13551885c99fb09994d285f7f7e87843fd2

Observation 1ade333e-63e3-43c6-9714-1cd9299d83c7 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Learning transferable visual models from natural language supervi- sion

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.188458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.188458Z digest=sha256:77c2130dfa2800a2052bd4daa18eda33a03679d1e7fb59d23ff38c7766ec9ce2

Observation d350a184-03ee-4255-a4d6-34504c349c54 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.235980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.235980Z digest=sha256:fc30b3b5bb7253695ab9019834477823b5e661f8cf14b1affb0e5ccc31102faf

Observation 0b906bca-8342-4e94-a1ca-9173e430b137 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.256485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.256485Z digest=sha256:6546fccec17cc375b08b5b257e24f78188bb879ecd37b123f7633e3ef2d00a76

Observation 11e5687e-4de8-4665-8bde-81fe11a06ab5 · outbound

This paper cites Linguistic bind- ing in diffusion models: Enhancing attribute correspondence through attention map alignment.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Linguistic bind- ing in diffusion models: Enhancing attribute correspondence through attention map alignment

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:55.870591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T15:08:54.270378Z digest=sha256:2add4cde66603ec354513a5112b7df25958cd1a96c6c89e77cc02abe50cf5ba3

Observation b3976570-40f6-4017-8709-f07fdd194195 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps High-resolution image synthesis with latent diffusion models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.331546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.331546Z digest=sha256:9099bd2303ed4f3ed64bd200be2e7e9b74d7911399f86cd7518e24b8b4954583

Observation e67aa0c6-9e1a-4956-a64a-a3190cd26d30 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Photorealistic text-to-image diffusion models with deep language understanding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:55.791771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T15:08:54.376477Z digest=sha256:9abc82012e9a1e939f876c36881b1e7523a6d5753f899ecf6a5550aa88571bed

Observation bfd31bc5-6cb1-49fd-be3d-344a8282fc4d · outbound

This paper cites Rethinking the spatial inconsistency in classifier- free diffusion guidance.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Rethinking the spatial inconsistency in classifier- free diffusion guidance

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:55.762366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T15:08:54.386550Z digest=sha256:a3c15904f0ead634d40b76681dad856fbaae85cca66ebcdd81aea2dc4464e750

Observation fd48f046-aa25-41ac-bf2b-ac475226fc33 · outbound

This paper cites Massive Activations in Large Language Models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Massive Activations in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.393582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.393582Z digest=sha256:c1c5dd87819d2e8562216f72e6602203bae13c7f4c4179ad8eb6c454183ac0b7

Observation f95914e2-e112-4fba-9254-4cc7a3ab33f7 · outbound

This paper cites Attention is all you need.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Attention is all you need

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.448164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.448164Z digest=sha256:0c0490d92964186f9011d5e0d637ddf6aa1c6337144810b71fd72dd25c0abe62

Observation 61e611f5-978f-42c8-af57-4d3effba7709 · outbound

This paper cites TokenCompose: Text-to-Image Diffusion with Token-level Supervision.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps TokenCompose: Text-to-Image Diffusion with Token-level Supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.481001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.481001Z digest=sha256:f34c95e1eed471e224c4690917c602b840b8eecff5ab45c2232d7ca00ffc0fb3

Observation 7ea65f26-5a4c-4102-a902-2ab30b770d93 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Efficient Streaming Language Models with Attention Sinks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.493758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.493758Z digest=sha256:f88989d3060f9b21e15a8ef086f39e8c44bc8dc7159c2123668c1233c2628718

Observation 68feadc4-3997-4a2a-89c4-757edabb1c2d · outbound

This paper cites Dynamic prompt learning: Addressing cross- attention leakage for text-based image editing.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Dynamic prompt learning: Addressing cross- attention leakage for text-based image editing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:55.688429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T15:08:54.508264Z digest=sha256:e334412bd08826cd948d4ce10e5125807a5ff671645ddbdf3558650e1e107bd4

Observation f24ac3fb-769c-4092-b47d-a0ea293bd3a1 · outbound

This paper cites Towards Understanding the Working Mechanism of Text-to-Image Diffusion Model.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Towards Understanding the Working Mechanism of Text-to-Image Diffusion Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.570225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.570225Z digest=sha256:75053e6ffd29c528d928cf6d059a77b4adbe725dd5e2921c4fb225667d803f4e

Observation 5a92f757-1784-42dd-8ef2-10fe272ceac7 · outbound

This paper cites Uncovering the Text Embedding in Text-to-Image Diffusion Models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Uncovering the Text Embedding in Text-to-Image Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.605688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.605688Z digest=sha256:edf06846b2ad923bc0bd30d5016d1cc7f5c16091b3eae8573aa39ed06034ef7a

Observation 988c4283-3a19-43a4-aa41-6e97888d8372 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.617831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.617831Z digest=sha256:e15b22878e5d884ec1b74a469d82dc90ffa349e06f0f84c4043cfea6497aabc4

Observation 1c5deeee-8208-4051-8fa1-6835cb49f8d4 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.627473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.627473Z digest=sha256:9c733f54112a57e23486be49fc011e4518f43df25575a032f47aaee43fab8cec

Observation 2d139588-cb0e-453b-8fe0-eca0c44a2ca8 · outbound

This paper cites Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.669287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.669287Z digest=sha256:d665e73b2edc0a05a7d7f7f628f96ab2d713de9a8295c142deb1aa19720e81d7

Pith citing papers

Observation fdbcc39c-bc22-4799-9ead-0a82d02e5fc4 · inbound

Detail++: Training-Free Detail Enhancer for T2I Diffusion Models cites this paper.

Detail++: Training-Free Detail Enhancer for T2I Diffusion Models Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:53.297589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:53.297589Z digest=sha256:82f0cd5a7e50d7ce6b3704a9468d63da09283668a11db94c02cb23110a4e7dd4

Observation 46a62a12-f3ce-40f9-a558-0886d5c4ec17 · inbound

DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation cites this paper.

DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:58:37.265067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-03T15:56:38.037304Z digest=sha256:d46da011e2a7908fd0c8a260bffb89de52a48fa34d2067b48037b8d858e84206