Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T09:25:19.066598Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2605.21611.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T09:25:19.066598Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f9ee4736-29c4-4eb7-a044-59007715992e · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Ming-Omni: A Unified Multimodal Model for Perception and Generation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b5095ec4-53ae-4cf5-945c-c4decadf84b6 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Nougat: Neural Optical Understanding for Academic Documents
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8ce526d6-a5da-497b-a961-57b2ea8a45b8 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation InstructPix2Pix: Learning to Follow Image Editing Instructions
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 28a45d0d-6e1c-47ce-88c4-d9bff06cbdb3 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Textdiffuser: Diffusion models as text painters
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b7b9a8ed-a566-4ab9-9b41-df3fb1a561a9 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Anydoor: Zero- shot object-level image customization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5d7c2be3-5219-4df0-a6e0-dd70cddb09a1 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8524f1f3-0e75-474c-b4f6-398f929e8a88 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Diffusion Models Beat GANs on Image Synthesis
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 83f599ec-a41a-4cfb-ac98-f26ecb5157bf · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Demystifying Flux Architecture
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9488ac53-35bc-469f-a3b8-3201efc5f28c · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Prompt-to-Prompt Image Editing with Cross Attention Control
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e4866763-e7c6-4584-aac1-374522c0023a · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Denoising Diffusion Probabilistic Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d7e8133b-e81c-4b8d-b985-1469f3022c78 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bd08cfa4-afd8-4379-8129-51fe1fea28b8 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation A style-based generator architecture for generative adversar- ial networks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2e377043-cd8f-49d2-9bab-2a21f7d1a85a · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Musiq: Multi-scale image quality transformer
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9b53c889-ee76-4457-9a18-170b68a99b45 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 653956ce-3acc-4992-8711-f570c1e5b97a · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Gligen: Open-set grounded text-to-image generation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 99ca2bf9-9662-4d98-8565-4c15972f8df9 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 37d2f4f1-d39b-437b-a7c6-82ea543d0408 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Repaint: Inpainting using denoising diffusion probabilistic models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c62f74dd-006a-4a95-800c-4c0ba9772ca1 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9aa96945-010b-4485-a1d9-8efec66cfa4e · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Semantic image synthesis with spatially-adaptive normalization
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2f98f065-2389-4ab9-9e81-2186815c78df · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Scalable diffusion models with transformers
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1c763e88-d7a7-4421-975b-499e995bb117 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Learning Transferable Visual Models From Natural Language Supervision
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d520d741-dfc8-46af-be69-95eb11f1ca9c · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 42a2e913-e131-413e-a98e-a300fd8a95da · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Scaling Up Models and Data with $\texttt{t5x}$ and $\texttt{seqio}$
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 354ddd2e-350f-488f-88f1-7f3289730960 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation High-Resolution Image Synthesis with Latent Diffusion Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9c25e2c3-6a69-4367-af4b-7e84cb1db5a1 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dc7f86de-d30c-4621-9a55-0870c79384f0 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Ominicontrol: Min- imal and universal control for diffusion transformer
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d3e474c7-4cb7-4026-9265-8ff4885844b6 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Omni-video: Democra- tizing unified video understanding and generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c5fc9805-f99b-445c-814b-aba84d2adc10 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation AnyText: Multilingual Visual Text Generation And Editing
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 53fcbf81-7705-44f2-bdd4-9d5f88fa8717 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d1d634a9-f8ad-49a2-a030-998f632dc134 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation DeepSeek-OCR: Contexts Optical Compression
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8c761652-e10c-45df-983f-fb3b13f4f356 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Deepseek-ocr 2: Visual causal flow
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b8705968-57a0-4b7e-b311-35719e7ce089 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8f3d3cf9-c942-44af-ad24-af88ed4a70f8 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Omnigen: Unified image generation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ab611996-7029-4139-8fd1-d9aa9b7c8fdb · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Paint by example: Exemplar-based image editing with diffusion models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fbfdd351-9b9c-4eae-a68f-3166f4fc26da · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation ImgEdit: A Unified Image Editing Dataset and Benchmark
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 915eae3d-299d-4543-ab68-ea78f422d9ed · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Adding Conditional Control to Text-to-Image Diffusion Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0677353e-6415-4072-a3fe-cf3a01ae01dd · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Unpaired image-to-image translation using cycle-consistent adversarial networks
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 307974bf-dfd1-48a8-81d6-67dfa5595704 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation A task is worth one word: Learning with task prompts for high-quality versatile image inpainting
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7bbe8cac-2642-4c61-a2d0-8a9a4f521ca5 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6ccd17de-a318-4717-b189-794b0442c827 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9726b270-0018-4994-b3af-627d55015f88 · outbound
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation with CLIP semantic instruction
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
No inbound Pith citation observations are available.