Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T22:18:55.933311Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2605.12271.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T22:18:55.933311Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
66 of 66 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 34372a72-9492-48a2-9de7-3c7ecf3d6d81 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Qwen-Image Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8ddc56e3-df4a-4764-a7ec-2845b749391a · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Qwen2.5-vl technical report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c3de5697-f8a9-4455-a33e-637c7b81704b · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c0655809-0596-45f5-9125-81d6d22ffd2b · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132–52152
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7378e38a-87a4-4805-80b1-215eec5cfcbb · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm HunyuanVideo 1.5 Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 830d1764-e7c6-4a65-acd0-4cb5b1e2c3ef · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4ae3052f-3767-4116-9c0e-49e9de16edf3 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm High-resolution image synthesis with latent diffusion models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bce2f1e7-7eae-423b-b613-2691ca366fca · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Photorealistic text-to-image diffusion models with deep language understanding.Advances in Neural Information Processing Systems, 35:36479–36494
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5c0f9423-8d57-41f3-9f0e-da797f306971 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ee4cd683-eab7-4157-b8d4-b41800d55caa · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Scalable diffusion models with transformers
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ea4052fb-d679-4ea4-81d7-cc42b56b2611 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Redefining Temporal Modeling in Video Diffusion: The Vectorized Timestep Approach
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a627daad-e388-49be-8111-03359cfb701a · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Make-A-Video: Text-to-Video Generation without Text-Video Data
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1b38ec72-1b49-48d4-b1a8-6f89b7271d48 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8c4543ab-1372-4bb9-99ff-a937481e82ca · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2af313ef-0c1c-4574-8f6f-333274982623 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Movie Gen: A Cast of Media Foundation Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1106ef4b-2dca-4fab-b580-21bb30ce0f6f · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm VideoCrafter1: Open Diffusion Models for High-Quality Video Generation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ab10d082-d293-410c-b093-1956c3539a4c · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Evalcrafter: Benchmarking and evaluating large video generation models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation da952127-2102-44db-975a-79b5132ffccd · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation db2e0968-22bd-4b00-9db4-a9934965324e · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Learning transferable visual models from natural language supervision
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7317b8e6-7035-4fe5-bc5a-421b8f8e8114 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a66426f2-968e-4120-bc3a-b62884e0c948 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Prompt-to-Prompt Image Editing with Cross Attention Control
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3c80d119-73ab-4418-8651-03266d4a9214 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Instructpix2pix: Learning to follow image editing instructions
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9fa9e7e8-9e24-4690-9c7c-343075a88130 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm TokenFlow: Consistent Diffusion Features for Consistent Video Editing
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6e5cccbd-516b-4e97-8fe9-1cc14bde0e01 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8b80377b-c3f2-4c25-b3f5-aa0dab08222b · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Adding conditional control to text-to-image diffusion models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 56dda7ce-9b5f-414b-ad04-eb23813c1d3b · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c0a9aa0d-07a0-45b2-ad29-506a5f6e8acd · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Gligen: Open-set grounded text-to-image generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0d733e11-866c-4d0d-8f41-c42ce17f3208 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation beb42f9b-670f-4520-8650-3bab9246490c · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 441cbaee-5121-4a8a-8106-311eafb89e34 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ce856a71-1e92-473b-85c9-6468c9890ce8 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Videocomposer: Compositional video synthesis with motion controllability.Advances in Neural Information Processing Systems, 36:7594–7611
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d138ad65-c1a6-47cd-b2f7-ef89e511c0a8 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Omnigen: Unified image generation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e9c6ffa5-2c1e-4d62-8035-7e5b067cd766 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e1414055-4be3-47e2-a770-7e9ccc0d408e · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Vace: All-in- one video creation and editing
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8485f045-4432-4f0e-bf4f-46231a86e6a2 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Character-aware models improve visual text rendering
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c0a3d646-a242-4507-9269-5e9824a1fefe · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Glyphcontrol: Glyph conditional control for visual text generation.Advances in Neural Information Processing Systems, 36:44050–44066
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 487b97b5-e6d1-4f80-8d2c-0c55a89cfa75 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm AnyText: Multilingual Visual Text Generation And Editing
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 54256d8a-cafc-405b-bf4c-02fdbcffc5ec · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Textdiffuser: Diffusion models as text painters.Advances in Neural Information Processing Systems, 36: 9353–9387
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d8a22524-cabd-4247-a72a-6d023d450c46 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Improving image generation with better captions
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6bc2af46-c394-469b-8513-9d76ce622299 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Visual prompting via image inpainting.Advances in neural information processing systems, 35: 25005–25017
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 93cdd04e-ea6d-48a5-8a71-65b1ecd6a392 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm In-context learning unlocked for diffusion models.Advances in Neural Information Processing Systems, 36:8542–8562
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 96d56013-3980-4fe6-b534-6e54e29ee87b · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Context diffusion: In-context aware image generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7037d149-7862-4e18-b52c-1789bd1d74a0 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Stable diffusion models are secretly good at visual in-context learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8639671c-cbe6-449e-84ec-23716b5e730f · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Visualcloze: A universal image generation framework via visual in-context learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6f1e17da-0ba9-4c1e-aef2-05d6eeffa5e4 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Realgeneral: Unifying visual generation via temporal in-context learning with video models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bc3ed293-0cb5-48ed-9497-a0172144441f · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm UNIC: Unified In-Context Video Editing
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e33e2eea-d06b-4c15-ba1d-632fae2e7cf2 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Emu3.5: Native Multimodal Models are World Learners
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3f9372ae-1500-49c6-ad5e-e7d2938c04e6 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Show-o2: Improved Native Unified Multimodal Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f85f81c0-aaea-41f1-9d38-5b08f5451eb8 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Emerging Properties in Unified Multimodal Pretraining
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f919d687-9f33-4bbc-91b7-f97d15e76730 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 674feb39-9b0f-4fad-ab02-bfa10190aa6f · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f460fd29-d601-4093-8370-1dcdf5c1ee66 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 50ac9039-5608-455a-9201-3a517fe4f566 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 203a4fda-74fe-4a2e-8b72-50a298f3f88a · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Emu3: Next-Token Prediction is All You Need
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 976438e6-5f3e-4571-b215-73b27d26f3fc · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 973ff730-7b0c-42e0-950f-1fefdc5ec241 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Scaling rectified flow trans- formers for high-resolution image synthesis
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 635ecd46-57b7-485b-815e-58226d205a13 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c0ded2af-0e29-4fe6-94ac-7a7a5fe1a3b3 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3b5808f0-1db0-4753-b66e-0bd699d46489 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Lumina-image 2.0: A unified and efficient image generative framework
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dcf13539-75b9-4424-9a25-900f1f510917 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 73fead70-c284-4fb2-b0e4-e7b2e61a92b8 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fcf0f092-8bf2-4fda-84ed-c51eeabdbd38 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm https://platform.openai.com/docs/ models/gpt-image-1
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 40ef2f1a-2b6e-4061-bc22-cacf32f600b5 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Seedream 3.0 Technical Report
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a588c5d5-47df-4ba9-ab1f-f46ff7366eba · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm https://platform.openai.com/docs/ models/gpt-image-2
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ca13dcc8-3372-4dde-a174-006b76214fb5 · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm https://seed.bytedance.com/en/seedream5_ 0_lite
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d3ba91d1-bdc0-461f-965b-3d00e4a002fd · outbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm minor imperfections only
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.