Pith. sign in

Paper Citation Record · LEDGER

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects

As of 17 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2411.18936.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18936 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:46:55.077302Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:48:07.360602Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:48:09.935086Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 867b2174-f76d-4873-900d-be8b5d3c69f9 · outbound

This paper cites GPT-4 Technical Report.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.852033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.852033Z digest=sha256:d2c93f623ccce27afcd905e1c00a428998c744f27cb5d52ebbd489ea2bbf732d

Observation 27941070-00c5-4cb4-a400-6cdb83350e1b · outbound

This paper cites A-star: Test-time attention segregation and retention for text-to-image synthesis.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects A-star: Test-time attention segregation and retention for text-to-image synthesis

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.667455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.856579Z digest=sha256:47a506cdaf1983d5ddbe2f244c1059cceff3299f1ce6b16382a771e591b7148a

Observation e0cdd8dd-c864-4308-9bd8-79dd30cb69bd · outbound

This paper cites Separate-and-enhance: Compo- sitional finetuning for text2image diffusion models.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Separate-and-enhance: Compo- sitional finetuning for text2image diffusion models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.656574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.860537Z digest=sha256:0dcfc1e88a98e3ac4eee7f57b836f152c54478178a531733c5f995700d126d86

Observation d86b3594-ba4f-4313-bf25-612c40810975 · outbound

This paper cites Improving image generation with better captions.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Improving image generation with better captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.864734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.864734Z digest=sha256:d9abfca762b07fd4325590e626134189fa8bbe37089c757373586b8edcbea97f

Observation 759e9d3b-a68b-44c3-b6ad-9b681a367431 · outbound

This paper cites Video generation models as world simulators.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Video generation models as world simulators

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.868259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.868259Z digest=sha256:8fd3b0ce163767911d328bcdf363d7ed1e8aa1d963622c802727d9cd2bdb11bb

Observation 63e19866-2722-445d-9fd5-6ecdb1430584 · outbound

This paper cites Masactrl: Tuning-free mu- tual self-attention control for consistent image synthesis and editing.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Masactrl: Tuning-free mu- tual self-attention control for consistent image synthesis and editing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.871683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.871683Z digest=sha256:6da1392cd9e73b47b959a1e0ab03fe9bf81e29e19c48f0ae2231e1e1e28f8f40

Observation 12541d63-38f5-4e89-88b2-dac1e68a4a9e · outbound

This paper cites Freeman, Michael Ru- binstein, Yuanzhen Li, and Dilip Krishnan.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Freeman, Michael Ru- binstein, Yuanzhen Li, and Dilip Krishnan

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.627125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.875117Z digest=sha256:b7f27497f9036ee321c5a38361e9c436039259f396749fde247b1463fb80ef4d

Observation e94a5149-afc3-4dab-a174-8cc33276a1f2 · outbound

This paper cites Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.616520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.878462Z digest=sha256:47c74bb0f6b934d20fbf18881e3389c959d7e94999fc87c621a3ff4e6eb8bd5f

Observation 0fcc0d7c-14be-4f2d-aaed-56f4a2dabbbc · outbound

This paper cites Diffusion models beat gans on image synthesis.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Diffusion models beat gans on image synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.881670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.881670Z digest=sha256:2c4ff25dd3ae37fcfb8c22c44d667651c610e725ec3cddd151ed82e1a34c6ecc

Observation 1470ac48-1ead-4def-a714-1e42fe523b5c · outbound

This paper cites Cogview: Mastering text-to-image generation via transformers.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Cogview: Mastering text-to-image generation via transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.885388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.885388Z digest=sha256:692eb3bee9031a528bc23635058b96c3d28a64d880ea90e79e95b38348b8943e

Observation 73113ef5-4ee5-423f-8111-852a966c16c3 · outbound

This paper cites Cogview2: Faster and better text-to-image generation via hierarchical transformers.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Cogview2: Faster and better text-to-image generation via hierarchical transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.889094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.889094Z digest=sha256:6aef2c4ce3b8c3eb263877d139a3dfc671ff3ef7a9e6aa9266edb9e9d3cd5162

Observation 864d09cc-ae67-4797-9c84-e90360443a79 · outbound

This paper cites Diffusion self-guidance for control- lable image generation.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Diffusion self-guidance for control- lable image generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.892657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.892657Z digest=sha256:e5e7ba3c28941d5566a7cffd1a0cdbc4a9abc1e1319ea26d366ab6775cc0bbaf

Observation c3d5f95c-4055-4340-844a-7f798fa8e773 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.580674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.896048Z digest=sha256:fe48ff465f041a52aaf4062578d4c3e311fcf9a5ec090e4f1648a9ddb11c6604

Observation 97093992-3930-40c8-ac6d-4bf0e379f2bb · outbound

This paper cites Make-a-scene: Scene- based text-to-image generation with human priors.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Make-a-scene: Scene- based text-to-image generation with human priors

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.899545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.899545Z digest=sha256:643e30cc372dcaaf4883a7069429b9516f14b1b583946e7847750c7325ee2672

Observation e6c6f8fa-2774-4b29-8ee4-4131d092d389 · outbound

This paper cites Expressive text-to-image generation with rich text.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Expressive text-to-image generation with rich text

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.564514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.902807Z digest=sha256:196139de72423d49e1c8cca8f691d371661d509439cf3047927c958686f462f6

Observation 33135ddf-8a47-4304-8cc9-ea051ab99c48 · outbound

This paper cites Generative adversarial nets.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Generative adversarial nets

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.906003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.906003Z digest=sha256:a4d05d99aa427b5f5596efda1e197b1c6cd5881cd3b3cb1853ef5bc4e0a5b0df

Observation 2637f5ab-5837-412d-95c0-86fcc92028ee · outbound

This paper cites Initno: Boosting text-to-image diffu- sion models via initial noise optimization.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Initno: Boosting text-to-image diffu- sion models via initial noise optimization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.549133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.909153Z digest=sha256:a37aacad96595fc4847cd9db2468378d566300144ec02ac358bd30645d770816

Observation addfdf03-a8ca-4ed7-aba0-a5ebc9e224e9 · outbound

This paper cites Prompt-to-prompt image editing with cross attention control.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Prompt-to-prompt image editing with cross attention control

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.538961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.912354Z digest=sha256:ada6c4e2caa465765347dc8d998ffe607ed402fac16eb8595d295026ca01ac1c

Observation 4ec12c56-1edf-4a05-b302-5d5560a40d7d · outbound

This paper cites Prompt-to-prompt image editing with cross-attention control.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Prompt-to-prompt image editing with cross-attention control

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.915620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.915620Z digest=sha256:ebfcc37558fa559dc83f8d4f2db3d314fa2434d4098b32fdc0a57bd9dbceaa0a

Observation f1ce83cd-b86e-4e33-bcd2-41701f94d255 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Classifier-Free Diffusion Guidance

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.918893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.918893Z digest=sha256:f36fb040e5985d24036a2f42e80377ae05ad6018f9cba2aff3d2c2cfbfa7fe8d

Observation 7e969cce-2b23-445a-b967-4db2721e3494 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Denoising dif- fusion probabilistic models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.922703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.922703Z digest=sha256:afcc628442047272f7f6656ba2989b653687dee54dd1105335ab799ba30dabcb

Observation 40de2e34-552b-4b38-85df-341f1acbe14b · outbound

This paper cites Tifa: Accu- rate and interpretable text-to-image faithfulness evaluation with question answering.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Tifa: Accu- rate and interpretable text-to-image faithfulness evaluation with question answering

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.516535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.926060Z digest=sha256:7954559a2f9530464822b6470f3f6678054687a2e5a57ccb4ca4310f1889842f

Observation 89063730-bdb9-4440-b5a9-83d0f8c9f11e · outbound

This paper cites T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.929428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.929428Z digest=sha256:6a10dc4c46858038d5e993799727fdefd29f3218749bab342a40a2af94b5722f

Observation 2604f282-a306-4708-b82f-196b4ad2193e · outbound

This paper cites T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to- Image Generation.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to- Image Generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.506458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.933157Z digest=sha256:2eaee62064404a73050d2554e26892b1ad4782a4a6f8a31786c887d1a2ff5451

Observation e60c0b41-c7df-4726-948a-2fcf08a0e323 · outbound

This paper cites Scal- ing up gans for text-to-image synthesis.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Scal- ing up gans for text-to-image synthesis

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.495853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.936601Z digest=sha256:41909a64b2395748682a25d68cfafdab979d6beeb0074a7a6361bd35cbfdba1f

Observation 8ef1d3ce-a9d9-4aa7-acf2-5a92555836c5 · outbound

This paper cites Dense text-to-image generation with attention modulation.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Dense text-to-image generation with attention modulation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.485504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.939778Z digest=sha256:97c9d99ca70f4fbf8193595162e645ee194e713feca9abeff25e2080abaae12a

Observation c5de3cdd-9f9e-448d-936a-e10f5f1708a7 · outbound

This paper cites Autoregressive image generation using residual quantization.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Autoregressive image generation using residual quantization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.474750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.943132Z digest=sha256:161728184399d7728274dabca724ef5f70173dffcf0545820025244aea0caa9b

Observation f15c35e0-66c8-4730-8193-b2f9dc0c2330 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.946616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.946616Z digest=sha256:7778138017a24dd92cae89f9c7bcaa33ffa14569d5e7a130570c6914f3a4c6c4

Observation ffa5ea6d-4571-4692-9946-82782b49205f · outbound

This paper cites Pseudo numerical methods for diffusion models on manifolds.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Pseudo numerical methods for diffusion models on manifolds

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.950514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.950514Z digest=sha256:ac248fa1dd9d2efb4ae813821d84c9ff4b28430d0740a898c184f7b103625753

Observation d17f3bd6-1a0e-49ba-a2df-69e5a10cb0af · outbound

This paper cites Sdedit: Guided image synthesis and editing with stochastic differential equa- tions.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Sdedit: Guided image synthesis and editing with stochastic differential equa- tions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.953603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.953603Z digest=sha256:f9c1cbdc57e67f8fcaa141c381fcb4f9c0b7c5c67362b8ee4249421c2a58b042

Observation b697ecea-e62b-4fe1-ba2e-cfbb0f5f2868 · outbound

This paper cites Conform: Contrast is all you need for high- fidelity text-to-image diffusion models.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Conform: Contrast is all you need for high- fidelity text-to-image diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.445195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.956899Z digest=sha256:3d9595b97a083ab8aa359183169c64c23454141810d02aaeed17d93862e52dc9

Observation adbbb3eb-e8a8-4f6e-bc14-fb9a2c73992d · outbound

This paper cites Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.960844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.960844Z digest=sha256:11d5b412d087ac50837f5f0415f3532b9305e88884087da8f5005db56158fa72

Observation 2ccf1564-2804-4f7b-a964-353676dc680f · outbound

This paper cites Improved denoising diffusion probabilistic models.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Improved denoising diffusion probabilistic models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.429331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.964336Z digest=sha256:675ec634ccb1eca772fdfda48f31f23a77c1a5e9126f57e8488efb95b8e966ce

Observation 1122305e-8c16-4e90-a773-f79f73911fd3 · outbound

This paper cites A threshold selection method from gray- level histograms.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects A threshold selection method from gray- level histograms

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.419108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.967610Z digest=sha256:4c2a7fbf3578ba96dafeb891bb514f9baefcea2f5ce2e274814ba6d62e4064b7

Observation 8d48ec44-9e93-478e-a16a-89c2ab7024ae · outbound

This paper cites Grounded text-to-image synthesis with attention refocusing.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Grounded text-to-image synthesis with attention refocusing

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.408479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.970808Z digest=sha256:cdf7333f2fea70e0c31b03395dee5c535a75063a16b259394b5dd4c49488a283

Observation b31c765b-a960-42f4-bdc8-af40d73bfd53 · outbound

This paper cites SDXL: Improving latent diffusion models for high-resolution image synthesis.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects SDXL: Improving latent diffusion models for high-resolution image synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.974626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.974626Z digest=sha256:77fbf7e4ecec14404ab6450d725fff8bb1bc2aed80a000b11fb5d7acbf45832c

Observation 5729aef0-7f5d-4dbb-a28c-526398b2d8ee · outbound

This paper cites Dreamfusion: Text-to-3d using 2d diffusion.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Dreamfusion: Text-to-3d using 2d diffusion

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.977860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.977860Z digest=sha256:a9c6698f9c70d490f8f7fd57b764f06c3d9c9cc8b5d1090570d23fac3a4d785d

Observation c8c9ae3c-326f-4e93-b74e-5807aa627bef · outbound

This paper cites Zero-shot text-to-image generation.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Zero-shot text-to-image generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.980882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.980882Z digest=sha256:97b5b075a6fa22fed5015890dee136ea0cddf1ffff4bd4ccfa269b70517fdd20

Observation b00ca42a-6a0a-4b9e-a626-981c2617d325 · outbound

This paper cites Hierarchical text-conditional image gener- ation with clip latents.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Hierarchical text-conditional image gener- ation with clip latents

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.377534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.984187Z digest=sha256:97ecce599a4acccc91ac2bfba2ff8bc5e9b070e538501642918a010b1cf9c567

Observation 1d8d80a9-dd33-455b-878f-9d0d1cbe56ca · outbound

This paper cites Linguistic bind- ing in diffusion models: Enhancing attribute correspondence through attention map alignment.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Linguistic bind- ing in diffusion models: Enhancing attribute correspondence through attention map alignment

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.366864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.987622Z digest=sha256:64674d75d0ad7407938c23233f19186d10e9f0168b6593b26e9e266036b7c544

Observation 66ce4b90-e1e4-462a-bed7-82c972521315 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects High-resolution image syn- thesis with latent diffusion models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.355969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:54.990856Z digest=sha256:8fe6b17d5adc61cbd09ec52e61194d6cfb2ac0fc90d005a08f17c671f473d19d

Observation be017625-1c41-4d5b-805a-20c273e8204b · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects U-net: Convolutional networks for biomedical image segmentation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.994355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.994355Z digest=sha256:6d004068a5f2b5843843639f4dee38548a9aaa0e1e15749c320375518da35dac

Observation 339bb67e-5ac4-4133-a2b6-34fd8f8bc010 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Photorealistic text-to-image diffusion models with deep language understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:54.997753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:54.997753Z digest=sha256:62ba198d4ba784e23a52fe54f7b57edd2629099dad983fc4e298105c5cdd1a3b

Observation 45203a6c-627c-4bf2-8787-d74d65110756 · outbound

This paper cites Stylegan-t: Unlocking the power of gans for fast large-scale text-to-image synthesis.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Stylegan-t: Unlocking the power of gans for fast large-scale text-to-image synthesis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:55.001605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:55.001605Z digest=sha256:3bfe316df74b0e24279255df6e72a48c0c0e07a31523ac154529f4c8838b19ca

Observation 3162f971-c04b-4ca2-b9b2-215140295e1f · outbound

This paper cites A picture is worth a thousand words: Principled recaptioning improves image generation, 2023.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects A picture is worth a thousand words: Principled recaptioning improves image generation, 2023

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.326214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:55.006091Z digest=sha256:9807c711d622d1a36a5796cf403d59fbcb0f5b4e345c855e49ae662b8f94de96

Observation 14a1f090-b85c-4050-98de-e1b7ff3ba304 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Deep unsupervised learning using nonequilibrium thermodynamics

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.315648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:55.009632Z digest=sha256:d8873d43bb33908295d8108e95c65015cf9e72218275af569715d7603671a9f7

Observation a2f10c0d-f0df-4a27-bb82-72cd7c183bd3 · outbound

This paper cites Score-based generative modeling through stochastic differential equa- tions.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Score-based generative modeling through stochastic differential equa- tions

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:55.012974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:55.012974Z digest=sha256:df85cd8df4fc6453db79f64b3b05a013c3d44bbd17987cc2f8748aaa98d7ca3f

Observation 751b3846-7ff2-4080-9473-7580e8c215c9 · outbound

This paper cites Df-gan: A simple and effec- tive baseline for text-to-image synthesis.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Df-gan: A simple and effec- tive baseline for text-to-image synthesis

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:55.016575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:55.016575Z digest=sha256:433f57ea6d938d73956f42ba0f9f394b18d9efcc5b80aa01f193526da37c4124

Observation 64cd234c-1b99-421e-83c6-3df2b3ec6c14 · outbound

This paper cites Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:55.020093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:55.020093Z digest=sha256:31bafb9cc200167fc3fd78929cebe17f872beb8c391164ee309c72a829a9c2e2

Observation a2b3a364-427f-4739-a6b3-7933c27d4b65 · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Plug-and-play diffusion features for text-driven image-to-image translation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:55.024146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:55.024146Z digest=sha256:e874dbffc2518425a174e4bc798439e73730eee90a49b5f48dacaf049bb646a0

Observation d8bf8eae-59d9-44e0-84c0-a260a89af162 · outbound

This paper cites Attention is all you need.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Attention is all you need

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:55.027913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:55.027913Z digest=sha256:aa47331dfeaacdc2ce9f4971e80cefba1d64f904bdaf6235a2ba11aad712ab98

Observation 5de08dc8-3cd3-45e3-a70d-52d98259ef1a · outbound

This paper cites Tokencompose: Text-to-image diffusion with token-level supervision.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Tokencompose: Text-to-image diffusion with token-level supervision

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.279782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:55.031610Z digest=sha256:59cd95ad4d2633be6c2d6c15a87aa91a159f8c597a7d772a6b764f59bc79a147

Observation 6d55f517-c3d2-4fb6-ba80-f56bfc9b6fe9 · outbound

This paper cites Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:55.035179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:55.035179Z digest=sha256:0e799d43c787223372a13101a088edd99b2cfed69f5fd67d0f69620857fe0e49

Observation c0dd81ff-00e6-43cb-bf9c-adf146e8a570 · outbound

This paper cites Attngan: Fine- grained text to image generation with attentional generative adversarial networks.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Attngan: Fine- grained text to image generation with attentional generative adversarial networks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:55.038461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:55.038461Z digest=sha256:6c4118893877e15242da4e0a22949a125a2b10fb078ab8ea091e9e4de1879f63

Observation 23274e42-8353-4953-9797-68d205c5656d · outbound

This paper cites Scaling autoregressive models for content-rich text-to-image generation.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Scaling autoregressive models for content-rich text-to-image generation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.255285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:55.042010Z digest=sha256:cc48bf0cec2eda98ff660c38d6f7110d00f99e3e2181a13cd8e1a25518ee78d5

Observation eac800dc-0087-4118-8c34-c436b9866f7a · outbound

This paper cites Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:55.045572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:55.045572Z digest=sha256:03717b131c8a962dcaa9e7befa2c72de7d73c9f8a1ca8efd43b9e73e461eae2b

Observation 161816cb-b77d-4461-8f47-bc53124469ff · outbound

This paper cites Attention calibration for disentangled text-to-image person- alization.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Attention calibration for disentangled text-to-image person- alization

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.245243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:55.049424Z digest=sha256:18ef6bf82eb8e6803b28cb3a520590086224b60f3560dd7c5c2b61f52da6b2e1

Observation 5f1d241a-6087-4577-a6fc-1a5c6ee886a9 · outbound

This paper cites CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:55.052802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:55.052802Z digest=sha256:2666057df62e5f059a280af6fdbb12b5757ff15f373c5181b2c52f96c2df43b6

Observation 43fc530e-705c-4f2e-9741-529b7f6663f9 · outbound

This paper cites bear” self-attn “elephant.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects bear” self-attn “elephant

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.234494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:55.056723Z digest=sha256:946e9ec33e9206d063cccd0f679873d0c6573600035e3f3dc47bd364716e6cd0

Observation 040e0cde-1ad8-49b2-822b-8ba9aeb844b5 · outbound

This paper cites an unresolved cited work.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:46:55.223056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:55.060646Z digest=sha256:4c52e34f9084f40218227ad0a88282b58a92951dd57979f7dfac0ff271e015d1

Observation 2d03c113-cba0-4959-ba7a-3637af4c3500 · outbound

This paper cites Ignore style, object size in comparison to its surroundings.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Ignore style, object size in comparison to its surroundings

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.212371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:55.064030Z digest=sha256:ec7e004b0bd5c4c70947f273ea23ca78be1ee9bd44f46b17cfe167dc1f3b2708

Observation 4cf3b6bc-2fe1-4549-a5fd-c9ac979d0428 · outbound

This paper cites an unresolved cited work.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:46:55.201986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:55.067434Z digest=sha256:9b460336dbfde426157b8c7925e3fd6a3186027e52dd568206a4726b17096e40

Observation 6d558a4e-5fca-4224-a069-213ad6af3d6c · outbound

This paper cites Ignore style, object size in comparison to its surroundings.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Ignore style, object size in comparison to its surroundings

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.191121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:55.070796Z digest=sha256:376491b4bfbefb59331f23ae04dedfc661122994ef6edc04277baf9a51869769

Observation 32618c0d-ebfc-487e-8d00-0f9046ae8bcd · outbound

This paper cites A appears.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects A appears

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.179677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:55.073983Z digest=sha256:6a624f32b0dad302932a71956d65c8046c3902f94b7a2ae2938b8dabd7c1b02c

Observation 97315f23-0792-4307-a0bf-4bebb18562af · outbound

This paper cites As shown in Tab.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects As shown in Tab

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:55.168132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:55.077302Z digest=sha256:28ae60bf0b3315827db7c72a6c4b08403ea32a99dc9f6b0171570895ee346038

Pith citing papers

Observation 8e26f502-36de-45d3-84c9-5f0333ae9fb3 · inbound

Performance Plateaus in Inference-Time Scaling for Text-to-Image Diffusion Without External Models cites this paper.

Performance Plateaus in Inference-Time Scaling for Text-to-Image Diffusion Without External Models Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:48:10.022788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:48:07.360602Z digest=sha256:be12c7b2afe7238c2a4fdc345984c101433688092436d5d36c88fc923ed42bb1