Pith. sign in

Paper Citation Record · LEDGER

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps

As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2411.15236.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15236 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:08:54.669287Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:45:53.297589Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T15:58:37.263440Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 860654df-32b1-470a-a7be-77ce9b923569 · outbound

This paper cites A-star: Test-time attention segregation and retention for text-to-image synthesis.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps A-star: Test-time attention segregation and retention for text-to-image synthesis

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.587696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:08:53.220813Z digest=sha256:c08e45ff551f401bf6db0f9e31df45c6275adb90747a111c281f6dae059fc44d

Observation 00464a84-ec3f-481f-9cba-3fb45053ed2d · outbound

This paper cites AlignIT: Enhancing Prompt Alignment in Customization of Text-to-Image Models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps AlignIT: Enhancing Prompt Alignment in Customization of Text-to-Image Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:08:55.655150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:08:53.416975Z digest=sha256:df933a580228ce2bb6b7450d971532a7d6be4d65c1835675aa4d6a59828e6f79

Observation f3a26ef5-c45c-459a-a889-f967ed3a35fd · outbound

This paper cites Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.510085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:08:53.538102Z digest=sha256:f3b32f7bd99abefd7d79ee491bcb8afda8fb5bebeab640cd5a4c890b830b7a40

Observation 86b56c1a-ed34-4e6f-a549-f9508e3e5072 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.586669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.586669Z digest=sha256:91ca7e8780f17d8d9a544cd59b49753f7224f3a0c7db20a58b1ae88419387aaf

Observation c1746190-6bc2-4565-81e2-59ffb069b285 · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.596085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.596085Z digest=sha256:261dbfec23a13966612d54ccb179220fa2990f126e5fa6e3e5b573f3e7e12fa3

Observation 338f1936-24c5-4ed3-94a4-fa7ee20f7b0b · outbound

This paper cites Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.459247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:08:53.629726Z digest=sha256:c19de340cf1f9c0d6b6d0bb0ca5eb2c6e2f4c6cac61996b0a982c8490d9b8151

Observation a614f833-0c21-4618-8e76-d663731cda2e · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.675405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.675405Z digest=sha256:2231656dcdb940f2b054e8967d2efc53c231ce1a10ecb56241780d4a8bdd4bf0

Observation 38e3d833-4cdc-45bd-8fd6-117feeac26d1 · outbound

This paper cites Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.683502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.683502Z digest=sha256:3a4b1eec2d9c55f164a98dacca8f259ced5c46cdf1a2b68446de706d6429268c

Observation 735bb36e-7c43-4f19-bb83-85f56785282f · outbound

This paper cites When Attention Sink Emerges in Language Models: An Empirical View.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps When Attention Sink Emerges in Language Models: An Empirical View

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.730797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.730797Z digest=sha256:85bc603e06fcadaa206d59c1121d25fadb35eee8ee90b289d538a3d4fbb66c2f

Observation 9f994c93-3122-47c8-bd26-2ca7ec81dc13 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.768421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.768421Z digest=sha256:9bceffa87455690c7d389fa93d126d234cdf4e3a45089dd86fb01b8761fd9ecf

Observation e0a854e5-19ca-4449-8159-8bf6b8d2fed6 · outbound

This paper cites spacy 2: Natural lan- guage understanding with bloom embeddings, convolutional neural networks and incremental parsing.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps spacy 2: Natural lan- guage understanding with bloom embeddings, convolutional neural networks and incremental parsing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.362672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:08:53.783216Z digest=sha256:0b6c41640ed80e888f5b1b08b6664f0f4f288bad87ae76077be810c587dd35b5

Observation ce474741-af47-465b-bc87-cbc72ca61bd7 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.821243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.821243Z digest=sha256:a171cf28329d9f28b391046222b5a95f58b3deef90c8b43fc72fd1c43e6a8d18

Observation 7ffb49e3-94ef-49ad-a930-5195b5240431 · outbound

This paper cites Tifa: Accu- rate and interpretable text-to-image faithfulness evaluation with question answering.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Tifa: Accu- rate and interpretable text-to-image faithfulness evaluation with question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.288975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:08:53.842397Z digest=sha256:57b92f83af8cd9a1d101bc0c4df3d3034fa3dc26bbe8b622196e18d6a68e0592

Observation 4139fcd9-6cf9-4359-9bca-762418452722 · outbound

This paper cites MC$^2$: Multi-concept Guidance for Customized Multi-concept Generation.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps MC$^2$: Multi-concept Guidance for Customized Multi-concept Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.854867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.854867Z digest=sha256:a6ce3a6ffb1c87d81df2b3f879d5d04a762bdefbca2e27cfd949e1372a8b456d

Observation 166d9dcb-4ec2-453d-b425-f61bd04fea16 · outbound

This paper cites Dense text-to-image generation with attention modulation.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Dense text-to-image generation with attention modulation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.879934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.879934Z digest=sha256:fe18ee16b31c5633fb1a79b69e56cb697d54ecf881578d578df7b952ef33c6ce

Observation 9be8d80a-c345-45a0-a1d6-eb7b3402a058 · outbound

This paper cites mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.954810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.954810Z digest=sha256:83106f3a38850a04a8fd812c9c9928f7daefb599f7fe1dd0ed213c4b184c2388

Observation 978e7c52-b3e2-417e-9378-8a35852e42b8 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.982683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.982683Z digest=sha256:22eba6bc65957605039814dddf9bfd99b94337bd68deb146ecf911252438b94c

Observation cd0d00d4-ea19-4ec4-9c31-f439499692a8 · outbound

This paper cites Microsoft coco: Common objects in context.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Microsoft coco: Common objects in context

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.993512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.993512Z digest=sha256:df64b1183fbcb738156f230ad8f851f9b7ad0f9c6a3e1f3f5a9fb2e902562095

Observation 37b1587b-5b1e-4954-a3ad-5feb3b425a7a · outbound

This paper cites Improving Text-to-Image Consistency via Automatic Prompt Optimization.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Improving Text-to-Image Consistency via Automatic Prompt Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.008711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.008711Z digest=sha256:8def05ec6e10819c3be295640fd34271532cc868e47c5def83c8d98e870292a3

Observation c62ebd91-f847-4e46-a36a-22593e22af9e · outbound

This paper cites Attention Overlap Is Responsible for The Entity Missing Problem in Text-to-image Diffusion Models!.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Attention Overlap Is Responsible for The Entity Missing Problem in Text-to-image Diffusion Models!

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:08:55.214378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:08:54.067179Z digest=sha256:3de44ca41962340f4b150918decf341bcf7a7c789f48ecb254bd148ae37be890

Observation 22456177-2924-4e5a-80e3-d21454ed916a · outbound

This paper cites Conform: Contrast is all you need for high- fidelity text-to-image diffusion models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Conform: Contrast is all you need for high- fidelity text-to-image diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.127872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:08:54.125524Z digest=sha256:8a1f9aa8563bc5c912381f103a254989383acf641314267bf2be25389a3679f8

Observation 70120ce3-9e5c-4932-97f2-7d689dec5a89 · outbound

This paper cites Openai gpt-3 api [gpt-3.5-turbo], 2024.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Openai gpt-3 api [gpt-3.5-turbo], 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.003417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:08:54.138455Z digest=sha256:b1e0e6d0828e7cbf76099600cd0e86b35da67997ef8de70207f4e9f9ed766ed3

Observation 1b51419f-9f9c-4f71-bb8d-f1216a926d55 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.150477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.150477Z digest=sha256:0c9c9c781bf4b91265313a8600b6baa00ce4c2ea35a53712f293bccfb28d6931

Observation f7c116de-be8d-4031-88e7-7e733ee9434a · outbound

This paper cites Not All Noises Are Created Equally:Diffusion Noise Selection and Optimization.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Not All Noises Are Created Equally:Diffusion Noise Selection and Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.173401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.173401Z digest=sha256:8faf18e273f74397ebc2f529d174e87bda33909450aa338965f1b28eaa7e8c96

Observation 1ade333e-63e3-43c6-9714-1cd9299d83c7 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Learning transferable visual models from natural language supervi- sion

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.188458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.188458Z digest=sha256:0e54480667df9e7be7a7d669ebfaa06eb16f068e44c0bff0d56a1eaba0387bf6

Observation d350a184-03ee-4255-a4d6-34504c349c54 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.235980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.235980Z digest=sha256:3856415a10204d899750ae47fc820208196a810b1579ffe093b5e3d089da8eca

Observation 0b906bca-8342-4e94-a1ca-9173e430b137 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.256485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.256485Z digest=sha256:3b99c9ae5ac8a736846bb9d13b050aab86c943b235b067126f3aeafe8e4f40e3

Observation 11e5687e-4de8-4665-8bde-81fe11a06ab5 · outbound

This paper cites Linguistic bind- ing in diffusion models: Enhancing attribute correspondence through attention map alignment.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Linguistic bind- ing in diffusion models: Enhancing attribute correspondence through attention map alignment

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:55.870591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:08:54.270378Z digest=sha256:a9fced32f7e1bd2068fbbf5aa111c862f40b9213e4833f4e50afa5f89fdb90b6

Observation b3976570-40f6-4017-8709-f07fdd194195 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps High-resolution image synthesis with latent diffusion models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.331546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.331546Z digest=sha256:ec252989fb17f07894349f734a6d22008eaabb90e0c70a1887bdf939628dee08

Observation e67aa0c6-9e1a-4956-a64a-a3190cd26d30 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Photorealistic text-to-image diffusion models with deep language understanding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:55.791771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:08:54.376477Z digest=sha256:ca168f8edfcfda062a56bc03de5632f8dd088d151da0d17578b7e8fafd956073

Observation bfd31bc5-6cb1-49fd-be3d-344a8282fc4d · outbound

This paper cites Rethinking the spatial inconsistency in classifier- free diffusion guidance.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Rethinking the spatial inconsistency in classifier- free diffusion guidance

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:55.762366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:08:54.386550Z digest=sha256:4c3a20a166e9df883c60b31082db753bc93427b5774d2e81f0396b4658d4e44e

Observation fd48f046-aa25-41ac-bf2b-ac475226fc33 · outbound

This paper cites Massive Activations in Large Language Models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Massive Activations in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.393582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.393582Z digest=sha256:2494aaa308bd84c06a371538847fb5d4b19c456d7188e58c7e50cb486ee41f51

Observation f95914e2-e112-4fba-9254-4cc7a3ab33f7 · outbound

This paper cites Attention is all you need.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Attention is all you need

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.448164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.448164Z digest=sha256:76a5f574c925caaed01c53fca3bd6e547cb8a636757bb0152ca53fc7986625d2

Observation 61e611f5-978f-42c8-af57-4d3effba7709 · outbound

This paper cites TokenCompose: Text-to-Image Diffusion with Token-level Supervision.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps TokenCompose: Text-to-Image Diffusion with Token-level Supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.481001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.481001Z digest=sha256:f9e9dba622c6428540dca451d9ceb43607506138210eb6064392c1450d2d502b

Observation 7ea65f26-5a4c-4102-a902-2ab30b770d93 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Efficient Streaming Language Models with Attention Sinks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.493758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.493758Z digest=sha256:f2ae8fb342573fb37fc1ab6dd40213d6225f4878b5d11df886fd832b4ff105dd

Observation 68feadc4-3997-4a2a-89c4-757edabb1c2d · outbound

This paper cites Dynamic prompt learning: Addressing cross- attention leakage for text-based image editing.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Dynamic prompt learning: Addressing cross- attention leakage for text-based image editing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:55.688429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:08:54.508264Z digest=sha256:38c30ec8b2a997a129582fcea5f5a0cbeb4dfd39deed69d23a1b2b494c21143e

Observation f24ac3fb-769c-4092-b47d-a0ea293bd3a1 · outbound

This paper cites Towards Understanding the Working Mechanism of Text-to-Image Diffusion Model.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Towards Understanding the Working Mechanism of Text-to-Image Diffusion Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.570225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.570225Z digest=sha256:18ff6ad6e94f7f6fc599a8673a53ac9809b1e15d046a3ef76ae276b42a00146a

Observation 5a92f757-1784-42dd-8ef2-10fe272ceac7 · outbound

This paper cites Uncovering the Text Embedding in Text-to-Image Diffusion Models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Uncovering the Text Embedding in Text-to-Image Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.605688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.605688Z digest=sha256:f24fcfc82bc639a43f7796f8c3dfda864f48e314d5fcee6cdadf36c44b9884c6

Observation 988c4283-3a19-43a4-aa41-6e97888d8372 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.617831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.617831Z digest=sha256:effd594fb72f0f34dc61d9ba566b17a9151bb92ee73e08539c5089da93f687f3

Observation 1c5deeee-8208-4051-8fa1-6835cb49f8d4 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.627473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.627473Z digest=sha256:d2d803cb1e544c66414cec9d01ee86e94b418a131f66de3ed36ebaa5725684c1

Observation 2d139588-cb0e-453b-8fe0-eca0c44a2ca8 · outbound

This paper cites Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.669287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.669287Z digest=sha256:aab69efc1c37e76bfe207fd1ee722d5a327507821ccb783606059ff8e72f3573

Pith citing papers

Observation fdbcc39c-bc22-4799-9ead-0a82d02e5fc4 · inbound

Detail++: Training-Free Detail Enhancer for T2I Diffusion Models cites this paper.

Detail++: Training-Free Detail Enhancer for T2I Diffusion Models Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:53.297589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:53.297589Z digest=sha256:2342b047738881928568e1f96da2175ee9fe1d7452f920faee5a8d01a838e0c3

Observation 46a62a12-f3ce-40f9-a558-0886d5c4ec17 · inbound

DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation cites this paper.

DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:58:37.265067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-03T15:56:38.037304Z digest=sha256:3de708b79382b7e367bb2b07167f038825cee842f482b94c7e66034d472b7d31