Pith. sign in

Paper Citation Record · LEDGER

Setting the Stage: Text-Driven Scene-Consistent Image Generation

As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2512.12598.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.12598 v3

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T17:40:25.779794Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact24
  • verified fuzzy28
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2844178c-3cee-4f70-a0fe-721f8fcd312a · outbound

This paper cites Blended latent diffusion.ACM transactions on graphics (TOG), 42 (4):1–11.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Blended latent diffusion.ACM transactions on graphics (TOG), 42 (4):1–11

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.765437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:cba72e86ea07ea796f5e0d5e6841656ac8dee34eab8392c2327217f6721c3065

Observation 3364098a-256e-482c-b447-f32ee2f8b51c · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

Setting the Stage: Text-Driven Scene-Consistent Image Generation In- structpix2pix: Learning to follow image editing instructions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.770393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:f7f80a9721727f3e775fd845d97328628ced0b3663cbd378e52ea57aee811af1

Observation 2803aa8c-fdd6-4a77-9472-ebf5efe062ec · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:44:17.322134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:d4457b2f07426284468aa6fc597b7741535c6d4dc825c191e4f16e810fd4e709

Observation c035275d-37b9-4f89-a01f-685f6ab1cd83 · outbound

This paper cites DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data.

Setting the Stage: Text-Driven Scene-Consistent Image Generation DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-27T02:04:59.423634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:366cf560fc8fbd8fdab72cb49d49f0a1a59101811039ee3bece9402a40906bf7

Observation f8dc622f-c2d0-4f23-91a1-f73d8d8d23ac · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Setting the Stage: Text-Driven Scene-Consistent Image Generation An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:44:17.254045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:1ef8ea4561e1ad7cf7e1724acdd63607076b4137f7607fa1ca1f56e4b78f08c6

Observation 639a2cd5-34ff-49e8-a34f-bd4f58569c6c · outbound

This paper cites Sample and Computation Redistribution for Efficient Face Detection.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Sample and Computation Redistribution for Efficient Face Detection

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:44:17.273074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:caf60211b69ea14bc95dac672c6cb087e1005386ce1a124b3d7b155eeff7e11a

Observation a1f1f424-73e5-49b9-b62e-c732a363e8b7 · outbound

This paper cites CameraCtrl: Enabling Camera Control for Text-to-Video Generation.

Setting the Stage: Text-Driven Scene-Consistent Image Generation CameraCtrl: Enabling Camera Control for Text-to-Video Generation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:44:17.230172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:ee5db0e6070a4e51ea295579efb0168b75324a92075d1973cd810ee6a6cf1d69

Observation e4860a6b-921b-4252-b950-e8b27b77849b · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Setting the Stage: Text-Driven Scene-Consistent Image Generation GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:44:17.309434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:ab9ce99ba38f2192b53d14d3c5cfc79998fda955489b894a5b823cc26691b1ae

Observation 590ac162-0ee6-480c-8f2f-68d4e0d7c64d · outbound

This paper cites Magicfight: Personalized martial arts combat video generation.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Magicfight: Personalized martial arts combat video generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.744291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:38169289e7128fbc9f7aac872d161ffc13c09a91dfe9edd59c78ee13ece67660

Observation 10b5969a-2a63-4ab0-9095-38db36b4dd4e · outbound

This paper cites Dual-schedule inver- sion: Training-and tuning-free inversion for real image edit- ing.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Dual-schedule inver- sion: Training-and tuning-free inversion for real image edit- ing

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.783642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:5e5a691d9176f0e46dab2f7a5cd77915b28248b7a023510c48d0c08726416655

Observation d5ca2794-f90d-48c7-bbc9-01ec5835e62d · outbound

This paper cites M4V: Multimodal Mamba for Efficient Text-to-Video Generation.

Setting the Stage: Text-Driven Scene-Consistent Image Generation M4V: Multimodal Mamba for Efficient Text-to-Video Generation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:44:17.287078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:e4a43017b70495af0f59dfdb1fbc62d973f2904398ec7c880210165da0e21457

Observation 0c31571f-3db4-4970-8ff4-da88bc5e68ee · outbound

This paper cites Gen- erative photography: a systematic, constructive approach.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Gen- erative photography: a systematic, constructive approach

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.769801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:3ae590d484a7af397c0a4f6fd2e9ca24e5cb8107517d843b165895efc7aede0c

Observation 19644532-0b24-476e-befa-e0403aa474a0 · outbound

This paper cites Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.788880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:bedc038cef10c67431ca2918f0c16ae78515e2975f5f2ccdb25f49d61d9f7504

Observation 8abcd290-c764-402f-ad01-c89cb4e6c805 · outbound

This paper cites interactive sto- rytelling.

Setting the Stage: Text-Driven Scene-Consistent Image Generation interactive sto- rytelling

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.786130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:8a44cac85e85eaf97207245bfb4778dcf55da5672289a3711af1eb5b18cc3a23

Observation 31881e81-69eb-4538-ae80-d1fe6afc964d · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Setting the Stage: Text-Driven Scene-Consistent Image Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:44:17.307598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:ca2b0d4aca682e14fe49f5ac9763fabe649e96007a7dd2b172d7725c98ec1ab7

Observation e7ddcf0b-6cd8-4564-b092-5fee30770c12 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Setting the Stage: Text-Driven Scene-Consistent Image Generation FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:44:17.300854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:ff017c270c5c6722040acdfbfd6b4a850470a384f24e145f5c52ce2747e9d1e8

Observation 6d38be05-6e83-4c50-b083-5ed8096805f6 · outbound

This paper cites Control-nerf: Editable feature volumes for scene rendering and manipulation.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Control-nerf: Editable feature volumes for scene rendering and manipulation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.783240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:072de97051c62578308ad230acf2a25c199744ddc117b056b3c795b5d853d3b1

Observation 98d4c47c-a654-4eb9-844b-086c226d4ac7 · outbound

This paper cites Mat: Mask-aware transformer for large hole im- age inpainting.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Mat: Mask-aware transformer for large hole im- age inpainting

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.741692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:3436b421232bc222909dd9b1e36179d34d9cc5d97e5aee5e20f205e1a5d0d04c

Observation aed01b9a-b1cd-45b9-a615-d6a59c1f116a · outbound

This paper cites Storygan: A sequential conditional gan for story vi- sualization.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Storygan: A sequential conditional gan for story vi- sualization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.772917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:1af778cad39f69e45c07de5b20941ffc60ef87be04abad48e9adb8c4754f60c7

Observation 3ea5290d-c31f-4f54-85b0-3073d548e776 · outbound

This paper cites Photomaker: Customizing re- alistic human photos via stacked id embedding.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Photomaker: Customizing re- alistic human photos via stacked id embedding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.729986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:7aabffaec8a8f83c25b6c54ae48033799b93d88779c18be7177aa89d0c9b494d

Observation 842264af-e7b6-4e5e-a05f-bf771f7ee1b6 · outbound

This paper cites Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:01:19.940440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:2657e0a556695eb6b01da6b7c1a1c79c2a63a21de4cba471fbe6901a664c2093

Observation 7097b422-ffb6-4bf8-882f-4556a9741dea · outbound

This paper cites Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.767051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:406624aa95cd2c0de7cc56c5b74e13d3195ee787ae1c1e7fe131c5208d3666f9

Observation 15c90e28-f2ba-4603-93d6-841558ab69de · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.776273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:c29a81a9a7cad6cb8a81203f0cd6965a5b013c7295920291f23397fea3182ec9

Observation 66f18dc7-0205-4019-abec-de79b85843e5 · outbound

This paper cites Story-adapter: A training-free iterative framework for long story visualization.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Story-adapter: A training-free iterative framework for long story visualization

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:44:17.238550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:a3c4620c1196ed86bb9eda9a334defe99bce87033bcaa31b790804f02a9ef768

Observation 8d868800-42b8-4844-ac60-22cd681e4935 · outbound

This paper cites Synthesizing coherent story with auto-regressive la- tent diffusion models.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Synthesizing coherent story with auto-regressive la- tent diffusion models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.740081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:a1efa7a3fb9830de0cd27ea25c927c47ee4e65cb0e0cbe5da251e4611c28e416

Observation b06f6723-4f7e-4727-a808-28db816a4670 · outbound

This paper cites Make-a-story: Visual memory conditioned consistent story generation.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Make-a-story: Visual memory conditioned consistent story generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.761412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:9d5f26168831e9bbece49d0043129b46742d6852a9f3a8e2decc96705f12c130

Observation 52fac083-52f7-4f69-9ea1-d2cdba25d432 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Setting the Stage: Text-Driven Scene-Consistent Image Generation SAM 2: Segment Anything in Images and Videos

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:44:17.296344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:33cb44a5ea1bbcfd804b9c223d8ea794e4f26dabb5e13e7a6e5690de37209295

Observation 0f488faf-3791-47e2-837f-bfcc08d49bdc · outbound

This paper cites Minima: Modality invariant im- age matching.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Minima: Modality invariant im- age matching

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.791206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:2feef6b0dc76ac50137205ae53e843b680567836325d83ebf154fc321cec8910

Observation 73bf747c-a7dd-4497-81a9-04ee421d4353 · outbound

This paper cites Seedream 4.0: Toward Next-generation Multimodal Image Generation.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Seedream 4.0: Toward Next-generation Multimodal Image Generation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:44:17.295648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:c177acdd052eeb2f3f64e254052293a9fc7f2e7642c26b02059bebaf5fa074d2

Observation 575b06ba-c52c-4ca8-8c2c-ee13f50d0be6 · outbound

This paper cites Univst: A unified framework for training-free localized video style transfer.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Univst: A unified framework for training-free localized video style transfer

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:44:17.313773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:c7202e5d11fd260259766c3156c1f5bcd5aefa8b74257feee3854f7e441b1677

Observation d1ecbc1f-a54f-40d6-876f-e69a19975c00 · outbound

This paper cites Scenedecorator: Towards scene-oriented story generation with scene planning and scene consistency.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Scenedecorator: Towards scene-oriented story generation with scene planning and scene consistency

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:44:17.263091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:b73f5cfd5843bffed761669ba3a07e525fe5a4bc9b2acc5dcc3bfbf6fa5bbed6

Observation 40e085fd-de90-4e5b-bb83-2f80c7408c6d · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:44:17.304699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:7759d59529f22fd14499c7583a2fa2af76e2f16d58656dd45cb297c179813ffc

Observation fcee3dba-e2a2-4070-bc13-31d83bf5ff90 · outbound

This paper cites Vistadream: Sampling multiview consistent images for single-view scene reconstruction.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Vistadream: Sampling multiview consistent images for single-view scene reconstruction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.781464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:0d55c6caea33e213a3d63ba180d94b089157d8343aebd39d9b07f510fae4dacc

Observation 4a6deffd-c08a-4aad-ada4-d769ce313d79 · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

Setting the Stage: Text-Driven Scene-Consistent Image Generation InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:44:17.303492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:a9cc3bafeab377272f2e1615c59c28f1bead6c58db200f5571564f1581979f6d

Observation 73a8ebd4-e4e9-4a54-8ae7-e6d9f5933f83 · outbound

This paper cites StyleAdapter: A Unified Stylized Image Generation Model.

Setting the Stage: Text-Driven Scene-Consistent Image Generation StyleAdapter: A Unified Stylized Image Generation Model

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:44:17.277647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:d20c79c46f6de08e162f9f0cb2dba894e3b5feffd0744a26432075dce5c85064

Observation 2582e4a0-46ba-4846-9ff8-f45a23341e86 · outbound

This paper cites Omniedit: Building image edit- ing generalist models through specialist supervision.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Omniedit: Building image edit- ing generalist models through specialist supervision

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.730162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:bfd79f9f88032ff4f5218c1b982059befc59682b95e34a1a93a90a0b88a484e6

Observation 4e5e305c-f6e1-4e1c-8dd9-bd9181758744 · outbound

This paper cites Qwen-Image Technical Report.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Qwen-Image Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:44:17.299328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:580437071cc01af3445ed7575d269aa2f2b13689f40a4a00bad9d5a59c8c9d2a

Observation bf524a49-a8da-44b0-959e-b453241e66c5 · outbound

This paper cites Dreamomni2: Multimodal instruction-based editing and generation.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Dreamomni2: Multimodal instruction-based editing and generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:44:17.272745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:c23ffe2d48186cdf31cd9bd731c3c101e31fdefa26c54bebb2294510813a8b84

Observation 4e845e5c-996a-4193-87fa-1711cdc680b8 · outbound

This paper cites Fastcomposer: Tuning-free multi- subject image generation with localized attention.Interna- tional Journal of Computer Vision, 133(3):1175–1194.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Fastcomposer: Tuning-free multi- subject image generation with localized attention.Interna- tional Journal of Computer Vision, 133(3):1175–1194

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.746047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:ca3378e377963eb0e934042e5585091ec0f8b19c432655fa02711eb2e286aa77

Observation 6f842eac-a479-4d7d-8772-569316e94ea3 · outbound

This paper cites Smartbrush: Text and shape guided object inpainting with diffusion model.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Smartbrush: Text and shape guided object inpainting with diffusion model

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.752896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:e94ba3206467da6872de373b35f7b3f4e8ea820ad735c00b89e0b5edb8eae273

Observation da4fe5ae-57e4-401c-9490-152adbc31aa9 · outbound

This paper cites Paint by example: Exemplar-based image editing with diffusion mod- els.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Paint by example: Exemplar-based image editing with diffusion mod- els

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.764646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:1af64757053cd8c2637478908005cde1296f33c560fcdeb972c91c695935a074

Observation a00ac5bc-9c3e-4466-94d2-42db3726b2ff · outbound

This paper cites Seed-story: Multi- modal long story generation with large language model.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Seed-story: Multi- modal long story generation with large language model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.778944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:15191783f7022b0e7497d2461842f8bb237dbebdf1dd81e27e55104649a00c93

Observation 2826e154-e4a0-46e7-a598-f91ff5931322 · outbound

This paper cites StyDeco: Unsupervised Style Transfer with Distilling Priors and Semantic Decoupling.

Setting the Stage: Text-Driven Scene-Consistent Image Generation StyDeco: Unsupervised Style Transfer with Distilling Priors and Semantic Decoupling

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:44:17.229086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:7d62f9f315a3661d2cd1c6457c40727edbe02567c9c57dafe24fc7c356468df4

Observation 8cf86e30-ea07-4320-ae70-47f65bc33c2d · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Setting the Stage: Text-Driven Scene-Consistent Image Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:44:17.258090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:47c8a408758d223a5f0c1a7304670f361d217784b5cd7b25ef61c0e1477523bd

Observation 44ccc016-c7b9-4a2e-8825-55f1bacf5356 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Setting the Stage: Text-Driven Scene-Consistent Image Generation IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:44:17.317961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:5b9203c521fa8d20d04960b4852b085b39dbd6c52aa16a62608954e6c4444f56

Observation fddcb322-f1ab-4db3-a656-b2b8bddb9432 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Adding conditional control to text-to-image diffusion models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.767669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:8dcf66e778bf2a46bf52c7cfd7ccc5a45026d1248fd00d067a4626483598d46b

Observation 7118e286-56f0-4687-930b-df42c32a035d · outbound

This paper cites Places: A 10 million image database for scene recognition.IEEE Transactions on Pattern Analy- sis and Machine Intelligence.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Places: A 10 million image database for scene recognition.IEEE Transactions on Pattern Analy- sis and Machine Intelligence

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.733231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:b6eb04d7a0d580d68bd4680b24a944c39ce19dfc079641c977cdcb1ab7a40d55

Observation 7fc60c39-ea15-4488-9f55-8cdf754ddc24 · outbound

This paper cites MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models.

Setting the Stage: Text-Driven Scene-Consistent Image Generation MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:44:17.258824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:74181bd94298f1e9f994ab244a299071bdd4244f36c0920710e5ee1b7e26ccc0

Observation 771a5c95-6867-4118-881b-8355346d50e2 · outbound

This paper cites Storydiffusion: Consistent self- attention for long-range image and video generation.Ad- vances in Neural Information Processing Systems, 37: 110315–110340.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Storydiffusion: Consistent self- attention for long-range image and video generation.Ad- vances in Neural Information Processing Systems, 37: 110315–110340

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.773338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:343d61ad88126d736594f862257def2c8301f43c6313f0d4ad15552d56892219

Observation 177b6957-b887-41a9-abc5-866496e535d9 · outbound

This paper cites Text-Image Alignment Metric Selection For evaluating text–image alignment, we compare two met- rics:CLIP-T[5] andGemini 2.5 Flash Text–Image Alignment (G2.5F-TIA)[3].

Setting the Stage: Text-Driven Scene-Consistent Image Generation Text-Image Alignment Metric Selection For evaluating text–image alignment, we compare two met- rics:CLIP-T[5] andGemini 2.5 Flash Text–Image Alignment (G2.5F-TIA)[3]

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.775927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:2e44b92e970022f32631a6f904243050bde7e8a18f743bb5345a5d0e192ce6d5

Observation bda63fbf-be93-4162-afb9-06fc53c6e0b8 · outbound

This paper cites an unresolved cited work.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-21T17:44:17.743029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:9a3a5915d1c5e97d08e5fb6eca6a1ac9131a2fc2e5fa37687e397b7e85f93f6c

Observation 09643a59-d257-4487-9166-14b10ed2c5c7 · outbound

This paper cites an unresolved cited work.

Setting the Stage: Text-Driven Scene-Consistent Image Generation Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-21T17:44:17.714839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:e37c16fe48cfdb15f729a8261f6043b0243b107894f4723792af87abc672826d

Observation 70691daa-4901-49e6-b932-948e2f3d5a96 · outbound

This paper cites We explicitly describe the Gemini 2.5 Flash prompts used for automatic scoring and the annotation interface shown to hu- man raters.

Setting the Stage: Text-Driven Scene-Consistent Image Generation We explicitly describe the Gemini 2.5 Flash prompts used for automatic scoring and the annotation interface shown to hu- man raters

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.712257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:ea556dfce93f194136b2e2f4db504f140e653a5d96b655620641b13d5aa08719

Observation 1ce9a379-2b50-4616-b02b-34c5e603b35d · outbound

This paper cites These videos were synthesized using the Kling image-to-video model, utilizing keyframes produced by our method.

Setting the Stage: Text-Driven Scene-Consistent Image Generation These videos were synthesized using the Kling image-to-video model, utilizing keyframes produced by our method

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T17:44:17.717715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:24b600ad3c6f392ce2182173c1b69a76b1908b397863b94399587377cdd0a5b2

Pith citing papers

No inbound Pith citation observations are available.