Pith. sign in

Paper Citation Record · LEDGER

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation

As of 11 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2605.21611.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.21611 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T09:25:19.066598Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact23
  • verified fuzzy15
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9ee4736-29c4-4eb7-a044-59007715992e · outbound

This paper cites Ming-Omni: A Unified Multimodal Model for Perception and Generation.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.784211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:42c5194d2734ba989e077224ac6f7e1227a4a5e972a986a10520af5e7d472dc7

Observation b5095ec4-53ae-4cf5-945c-c4decadf84b6 · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Nougat: Neural Optical Understanding for Academic Documents

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.807752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:290fdec50c45bdd6a75874b581015679885d6ffd0f0e3a582840928043eb8b55

Observation 8ce526d6-a5da-497b-a961-57b2ea8a45b8 · outbound

This paper cites InstructPix2Pix: Learning to Follow Image Editing Instructions.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.779221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:07b67ea4150ef6f5fec6cdd16e036047c056ede09e40853cf3faa17cb1fccf7d

Observation 28a45d0d-6e1c-47ce-88c4-d9bff06cbdb3 · outbound

This paper cites Textdiffuser: Diffusion models as text painters.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Textdiffuser: Diffusion models as text painters

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.105030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:f4260b02d0cceb525426159724828fb41f3c7a8cdd44dfb3a63d1cefa780f97c

Observation b7b9a8ed-a566-4ab9-9b41-df3fb1a561a9 · outbound

This paper cites Anydoor: Zero- shot object-level image customization.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Anydoor: Zero- shot object-level image customization

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.101857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:2f418fff048ba3a2dcf45df8eb783cc7b677d7265cec4619c7489ef8b15ebaec

Observation 5d7c2be3-5219-4df0-a6e0-dd70cddb09a1 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.789280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:6031295fcdf0f92619779f83d4a9c663c787a89a386286ff1e86353fdba750f8

Observation 8524f1f3-0e75-474c-b4f6-398f929e8a88 · outbound

This paper cites Diffusion Models Beat GANs on Image Synthesis.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Diffusion Models Beat GANs on Image Synthesis

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.793708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:843777b37a1ff90ad72f875843851f95b8d5a28ce81af29f0fc9b26410ba9393

Observation 83f599ec-a41a-4cfb-ac98-f26ecb5157bf · outbound

This paper cites Demystifying Flux Architecture.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Demystifying Flux Architecture

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.803649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:1fafb7f91a5dadf38e25e1c2a06aaade6442367a4583786adfe826126f52ab51

Observation 9488ac53-35bc-469f-a3b8-3201efc5f28c · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.798428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:65a3f686fd380f652eafe3a21f0bfdd4cbc5264aba8e602b7f1bb9f1024d7cf2

Observation e4866763-e7c6-4584-aac1-374522c0023a · outbound

This paper cites Denoising Diffusion Probabilistic Models.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Denoising Diffusion Probabilistic Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.773573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:fb3252777dd75fa97590eb13ffc0e481e4f9887eac14592cae31cff417d14148

Observation d7e8133b-e81c-4b8d-b985-1469f3022c78 · outbound

This paper cites Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.108241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:57e68a6649b1fa92e2a2b2144225ecd7b7ff3ebe046874df2db17e2ab8a92f3f

Observation bd08cfa4-afd8-4379-8129-51fe1fea28b8 · outbound

This paper cites A style-based generator architecture for generative adversar- ial networks.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation A style-based generator architecture for generative adversar- ial networks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.078006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:b1dfcac6acdc644b61869c72f491101ddcc0910ad1c5c900304ec6cc817c9462

Observation 2e377043-cd8f-49d2-9bab-2a21f7d1a85a · outbound

This paper cites Musiq: Multi-scale image quality transformer.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Musiq: Multi-scale image quality transformer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.091733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:99955c8f8a5e12188c90a9a2811809a25b90da051f3f35728dc7f5c1cbcd29b2

Observation 9b53c889-ee76-4457-9a18-170b68a99b45 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.727279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:d66c17984b04b8ecabe75e43a4f881bdd97febe68932c83cbb5297b779cd6d9c

Observation 653956ce-3acc-4992-8711-f570c1e5b97a · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Gligen: Open-set grounded text-to-image generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.074980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:c6638730ee6a5284c188d36f3c88cd8947115d9f6bbf691c61aa0044180c340c

Observation 99ca2bf9-9662-4d98-8565-4c15972f8df9 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.722625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:8ef2def2c4aec0b75eb27603f15978bf42da6f1b68d21b1d15ff46c794de1065

Observation 37d2f4f1-d39b-437b-a7c6-82ea543d0408 · outbound

This paper cites Repaint: Inpainting using denoising diffusion probabilistic models.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Repaint: Inpainting using denoising diffusion probabilistic models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.081507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:18ae317decaecd49fe5e7b583b47e63514c393ae09c649d9ace4012db3140d77

Observation c62f74dd-006a-4a95-800c-4c0ba9772ca1 · outbound

This paper cites GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.712176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:5e2261ae8a9c28b9511fdfa9e5030cfca7d106af08a524d878884ae053e07753

Observation 9aa96945-010b-4485-a1d9-8efec66cfa4e · outbound

This paper cites Semantic image synthesis with spatially-adaptive normalization.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Semantic image synthesis with spatially-adaptive normalization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.064550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:782a75a22e59a0af5ee0551d8a55b7170045e1ebbfd2238898c956be4f3fbd8b

Observation 2f98f065-2389-4ab9-9e81-2186815c78df · outbound

This paper cites Scalable diffusion models with transformers.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Scalable diffusion models with transformers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.067867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:0995cbd3cce3743e20aa889029179ca87f6dc22299c78d4ff34929b828527ff8

Observation 1c763e88-d7a7-4421-975b-499e995bb117 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Learning Transferable Visual Models From Natural Language Supervision

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.751151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:de1d44d6358bcabf56842334c765335e9c19799ee7e6feba092b3b3b33e5ae22

Observation d520d741-dfc8-46af-be69-95eb11f1ca9c · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.759730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:51cdf561334c9245f9b1dad2898485365046435bc6af99ca3e2f1101f5030b50

Observation 42a2e913-e131-413e-a98e-a300fd8a95da · outbound

This paper cites Scaling Up Models and Data with $\texttt{t5x}$ and $\texttt{seqio}$.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Scaling Up Models and Data with $\texttt{t5x}$ and $\texttt{seqio}$

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.769472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:86752c2a6d5fda5a0277ded961bd46ce0914f02c98d637044c87da78221ab6af

Observation 354ddd2e-350f-488f-88f1-7f3289730960 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation High-Resolution Image Synthesis with Latent Diffusion Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.755259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:60dd6fb7051deae7138e041fa2fc34ad190dccdb1e9c516d0ef9eeb24211295d

Observation 9c25e2c3-6a69-4367-af4b-7e84cb1db5a1 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.071652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:fa55fcfc72fcd9397de379a07609a25d270ec11ece2d99ad82da339919d9025f

Observation dc7f86de-d30c-4621-9a55-0870c79384f0 · outbound

This paper cites Ominicontrol: Min- imal and universal control for diffusion transformer.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Ominicontrol: Min- imal and universal control for diffusion transformer

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.088603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:b4d6332d23bdb03f2e730b79e78e8e896974d816c726fb4ea7af389b8fd20cfc

Observation d3e474c7-4cb7-4026-9265-8ff4885844b6 · outbound

This paper cites Omni-video: Democra- tizing unified video understanding and generation.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Omni-video: Democra- tizing unified video understanding and generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.764404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:5a644d5409e5c98a0f7beefba278f54c509af3b84574454c22532e18aa0d5869

Observation c5fc9805-f99b-445c-814b-aba84d2adc10 · outbound

This paper cites AnyText: Multilingual Visual Text Generation And Editing.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation AnyText: Multilingual Visual Text Generation And Editing

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.717505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:5f9ba2ecb1919a980ad558b2579d597a1e9b73d126651d20c68d981d8a4e8269

Observation 53fcbf81-7705-44f2-bdd4-9d5f88fa8717 · outbound

This paper cites General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.695715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:db81334d309b9b9c238113758912ba899c4c00747d5b7b279925cdffe82e01fa

Observation d1d634a9-f8ad-49a2-a030-998f632dc134 · outbound

This paper cites DeepSeek-OCR: Contexts Optical Compression.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation DeepSeek-OCR: Contexts Optical Compression

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.706394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:a94e749e3f1aaa6241e1f531b45a0793125bf9916d5d1d5507513d2a24ad06c7

Observation 8c761652-e10c-45df-983f-fb3b13f4f356 · outbound

This paper cites Deepseek-ocr 2: Visual causal flow.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Deepseek-ocr 2: Visual causal flow

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.701364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:9718b35001c3176a838ce6e7db93aa166c08d8eafec6be48a6e3e05a54664c17

Observation b8705968-57a0-4b7e-b311-35719e7ce089 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.737616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:42517ccacaa6c7a26ec4eb3949705c2ff20fed1729d3225ae8c37c0e59815b20

Observation 8f3d3cf9-c942-44af-ad24-af88ed4a70f8 · outbound

This paper cites Omnigen: Unified image generation.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Omnigen: Unified image generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.085223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:232223a2f46ffb9246ef20a327a3dde3d5cbfbea25e3ae40b64594e8cecb1462

Observation ab611996-7029-4139-8fd1-d9aa9b7c8fdb · outbound

This paper cites Paint by example: Exemplar-based image editing with diffusion models.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Paint by example: Exemplar-based image editing with diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.095108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:acb8255d8f52b6aec06a17535f28057ba5783af41b749e53a7affdafd3f6b936

Observation fbfdd351-9b9c-4eae-a68f-3166f4fc26da · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.741995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:106c44433d93cf2d87fbe4375bc61c455bf4d7abc5a8b11b48269aa8524141ed

Observation 915eae3d-299d-4543-ab68-ea78f422d9ed · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Adding Conditional Control to Text-to-Image Diffusion Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.746997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:cd579a36093c8cd8763dce4a84b7cd9fbe7f112c472fe7e71380c86d60ed2139

Observation 0677353e-6415-4072-a3fe-cf3a01ae01dd · outbound

This paper cites Unpaired image-to-image translation using cycle-consistent adversarial networks.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Unpaired image-to-image translation using cycle-consistent adversarial networks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.098748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:22ef750a480b8b73e603c77f90319f30653ada3f5f0f1326ec23e3d521e679fd

Observation 307974bf-dfd1-48a8-81d6-67dfa5595704 · outbound

This paper cites A task is worth one word: Learning with task prompts for high-quality versatile image inpainting.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation A task is worth one word: Learning with task prompts for high-quality versatile image inpainting

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.061206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:66087c1a94fdf21f966c95ffd103d4c0103a621e957018545c0123bb0f6f6476

Observation 7bbe8cac-2642-4c61-a2d0-8a9a4f521ca5 · outbound

This paper cites an unresolved cited work.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-22T09:26:21.054117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:630d578bfe60b8f1dd103daa15184a0d793c32e3e7287762280f2108ed6bed38

Observation 6ccd17de-a318-4717-b189-794b0442c827 · outbound

This paper cites an unresolved cited work.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-22T09:26:21.057550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:66368837969cf077808ea600300cd7aafff717c3be641d22ed08fd22c42faf0e

Observation 9726b270-0018-4994-b3af-627d55015f88 · outbound

This paper cites with CLIP semantic instruction.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation with CLIP semantic instruction

Reference 42

Resolution
malformed identifier
arxiv_id, observed 2026-05-22T09:26:20.732509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:5000d4942914f723b584d6c35885c425a16c4a393154bef37999f0dce8b553e9

Pith citing papers

No inbound Pith citation observations are available.