Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T17:40:25.779794Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2512.12598.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T17:40:25.779794Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2844178c-3cee-4f70-a0fe-721f8fcd312a · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Blended latent diffusion.ACM transactions on graphics (TOG), 42 (4):1–11
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3364098a-256e-482c-b447-f32ee2f8b51c · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation In- structpix2pix: Learning to follow image editing instructions
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2803aa8c-fdd6-4a77-9472-ebf5efe062ec · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c035275d-37b9-4f89-a01f-685f6ab1cd83 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f8dc622f-c2d0-4f23-91a1-f73d8d8d23ac · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 639a2cd5-34ff-49e8-a34f-bd4f58569c6c · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Sample and Computation Redistribution for Efficient Face Detection
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a1f1f424-73e5-49b9-b62e-c732a363e8b7 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation CameraCtrl: Enabling Camera Control for Text-to-Video Generation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e4860a6b-921b-4252-b950-e8b27b77849b · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 590ac162-0ee6-480c-8f2f-68d4e0d7c64d · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Magicfight: Personalized martial arts combat video generation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 10b5969a-2a63-4ab0-9095-38db36b4dd4e · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Dual-schedule inver- sion: Training-and tuning-free inversion for real image edit- ing
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d5ca2794-f90d-48c7-bbc9-01ec5835e62d · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation M4V: Multimodal Mamba for Efficient Text-to-Video Generation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0c31571f-3db4-4970-8ff4-da88bc5e68ee · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Gen- erative photography: a systematic, constructive approach
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 19644532-0b24-476e-befa-e0403aa474a0 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8abcd290-c764-402f-ad01-c89cb4e6c805 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation interactive sto- rytelling
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 31881e81-69eb-4538-ae80-d1fe6afc964d · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e7ddcf0b-6cd8-4564-b092-5fee30770c12 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6d38be05-6e83-4c50-b083-5ed8096805f6 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Control-nerf: Editable feature volumes for scene rendering and manipulation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 98d4c47c-a654-4eb9-844b-086c226d4ac7 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Mat: Mask-aware transformer for large hole im- age inpainting
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aed01b9a-b1cd-45b9-a615-d6a59c1f116a · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Storygan: A sequential conditional gan for story vi- sualization
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3ea5290d-c31f-4f54-85b0-3073d548e776 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Photomaker: Customizing re- alistic human photos via stacked id embedding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 842264af-e7b6-4e5e-a05f-bf771f7ee1b6 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7097b422-ffb6-4bf8-882f-4556a9741dea · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 15c90e28-f2ba-4603-93d6-841558ab69de · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 66f18dc7-0205-4019-abec-de79b85843e5 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Story-adapter: A training-free iterative framework for long story visualization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8d868800-42b8-4844-ac60-22cd681e4935 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Synthesizing coherent story with auto-regressive la- tent diffusion models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b06f6723-4f7e-4727-a808-28db816a4670 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Make-a-story: Visual memory conditioned consistent story generation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 52fac083-52f7-4f69-9ea1-d2cdba25d432 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation SAM 2: Segment Anything in Images and Videos
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0f488faf-3791-47e2-837f-bfcc08d49bdc · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Minima: Modality invariant im- age matching
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 73bf747c-a7dd-4497-81a9-04ee421d4353 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Seedream 4.0: Toward Next-generation Multimodal Image Generation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 575b06ba-c52c-4ca8-8c2c-ee13f50d0be6 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Univst: A unified framework for training-free localized video style transfer
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d1ecbc1f-a54f-40d6-876f-e69a19975c00 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Scenedecorator: Towards scene-oriented story generation with scene planning and scene consistency
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 40e085fd-de90-4e5b-bb83-2f80c7408c6d · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Wan: Open and Advanced Large-Scale Video Generative Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fcee3dba-e2a2-4070-bc13-31d83bf5ff90 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Vistadream: Sampling multiview consistent images for single-view scene reconstruction
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4a6deffd-c08a-4aad-ada4-d769ce313d79 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation InstantID: Zero-shot Identity-Preserving Generation in Seconds
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 73a8ebd4-e4e9-4a54-8ae7-e6d9f5933f83 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation StyleAdapter: A Unified Stylized Image Generation Model
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2582e4a0-46ba-4846-9ff8-f45a23341e86 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Omniedit: Building image edit- ing generalist models through specialist supervision
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4e5e305c-f6e1-4e1c-8dd9-bd9181758744 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Qwen-Image Technical Report
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bf524a49-a8da-44b0-959e-b453241e66c5 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Dreamomni2: Multimodal instruction-based editing and generation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4e845e5c-996a-4193-87fa-1711cdc680b8 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Fastcomposer: Tuning-free multi- subject image generation with localized attention.Interna- tional Journal of Computer Vision, 133(3):1175–1194
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6f842eac-a479-4d7d-8772-569316e94ea3 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Smartbrush: Text and shape guided object inpainting with diffusion model
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation da4fe5ae-57e4-401c-9490-152adbc31aa9 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Paint by example: Exemplar-based image editing with diffusion mod- els
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a00ac5bc-9c3e-4466-94d2-42db3726b2ff · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Seed-story: Multi- modal long story generation with large language model
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2826e154-e4a0-46e7-a598-f91ff5931322 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation StyDeco: Unsupervised Style Transfer with Distilling Priors and Semantic Decoupling
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8cf86e30-ea07-4320-ae70-47f65bc33c2d · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 44ccc016-c7b9-4a2e-8825-55f1bacf5356 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fddcb322-f1ab-4db3-a656-b2b8bddb9432 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Adding conditional control to text-to-image diffusion models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7118e286-56f0-4687-930b-df42c32a035d · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Places: A 10 million image database for scene recognition.IEEE Transactions on Pattern Analy- sis and Machine Intelligence
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7fc60c39-ea15-4488-9f55-8cdf754ddc24 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 771a5c95-6867-4118-881b-8355346d50e2 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Storydiffusion: Consistent self- attention for long-range image and video generation.Ad- vances in Neural Information Processing Systems, 37: 110315–110340
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 177b6957-b887-41a9-abc5-866496e535d9 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Text-Image Alignment Metric Selection For evaluating text–image alignment, we compare two met- rics:CLIP-T[5] andGemini 2.5 Flash Text–Image Alignment (G2.5F-TIA)[3]
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bda63fbf-be93-4162-afb9-06fc53c6e0b8 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 09643a59-d257-4487-9166-14b10ed2c5c7 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 70691daa-4901-49e6-b932-948e2f3d5a96 · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation We explicitly describe the Gemini 2.5 Flash prompts used for automatic scoring and the annotation interface shown to hu- man raters
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1ce9a379-2b50-4616-b02b-34c5e603b35d · outbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation These videos were synthesized using the Kling image-to-video model, utilizing keyframes produced by our method
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.