Pith. sign in

Paper Citation Record · LEDGER

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm

As of 6 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2605.12271.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12271 v2

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T22:18:55.933311Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact31
  • verified fuzzy33
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 34372a72-9492-48a2-9de7-3c7ecf3d6d81 · outbound

This paper cites Qwen-Image Technical Report.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Qwen-Image Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.680500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:5e386b437c1a41ebeab08715f9c815cd6dcf545d0fe4736c4c80d5f313425aa2

Observation 8ddc56e3-df4a-4764-a7ec-2845b749391a · outbound

This paper cites Qwen2.5-vl technical report.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Qwen2.5-vl technical report

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.538628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:39c94c294c024aed4458408cfe7e4c7ad906fd0f42b5813568d5bc9aa3b885f5

Observation c3de5697-f8a9-4455-a33e-637c7b81704b · outbound

This paper cites Qwen2.5-VL Technical Report.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.673073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:7baaf475345250fcc2dfa7e832aab427281ae87c5e2c6acdd70536bdc71588c9

Observation c0655809-0596-45f5-9125-81d6d22ffd2b · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132–52152.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132–52152

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.540684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:e790141112239a228bc82b9a558b601d2100dcdf7f32467198fe9a1bc96d2946

Observation 7378e38a-87a4-4805-80b1-215eec5cfcbb · outbound

This paper cites HunyuanVideo 1.5 Technical Report.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm HunyuanVideo 1.5 Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.675275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:c7a443bf00bb041d4c7d835a746feb2401b37ebb15f10a30d620eab7a23c7c4c

Observation 830d1764-e7c6-4a65-acd0-4cb5b1e2c3ef · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.527142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:1b5bed9c9b4fe527ad1cf15f6bc60a02bcce5fc0dc4d6926ed0eff1c5f23b45c

Observation 4ae3052f-3767-4116-9c0e-49e9de16edf3 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm High-resolution image synthesis with latent diffusion models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.528937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:796bfbcff5845ae0cbc63efbc564438421a5caf570ca3faed39cfe6e98dd314d

Observation bce2f1e7-7eae-423b-b613-2691ca366fca · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.Advances in Neural Information Processing Systems, 35:36479–36494.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Photorealistic text-to-image diffusion models with deep language understanding.Advances in Neural Information Processing Systems, 35:36479–36494

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.533120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:b6fcf346f36b953a31a69748b0ac00ce76ef23e8ad5922f7b5e197b2862e715a

Observation 5c0f9423-8d57-41f3-9f0e-da797f306971 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.678279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:de6918b4230ec2024c4f773c3967c89d70a7103485def706c5fc9a025ff7f586

Observation ee4cd683-eab7-4157-b8d4-b41800d55caa · outbound

This paper cites Scalable diffusion models with transformers.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Scalable diffusion models with transformers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.530839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:98162e6862688deecc16d7b1ee22447d62cc2d747d51e264bc6d0c27e4acff1c

Observation ea4052fb-d679-4ea4-81d7-cc42b56b2611 · outbound

This paper cites Redefining Temporal Modeling in Video Diffusion: The Vectorized Timestep Approach.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Redefining Temporal Modeling in Video Diffusion: The Vectorized Timestep Approach

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:46.683012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:a66fd0f6374080f91613cd88b8b1cb6f3e7e09dc37829e6819921b3ac32fba45

Observation a627daad-e388-49be-8111-03359cfb701a · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.663509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:b70884a8d2dd465f949e088c884fa1b0b857513ef2b3714cfe87e7280accae4e

Observation 1b38ec72-1b49-48d4-b1a8-6f89b7271d48 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.661038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:03b2f9d6b842683e6228513fec7baac5cf725863abbf53705692c14fb99434a2

Observation 8c4543ab-1372-4bb9-99ff-a937481e82ca · outbound

This paper cites an unresolved cited work.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-07-07T14:13:49.536757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:30193392771f35dcc478b3b98730bc4f91e9dde842e981162daf3146757487b3

Observation 2af313ef-0c1c-4574-8f6f-333274982623 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Movie Gen: A Cast of Media Foundation Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.594464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:98710a307ba6c4f6bbbd46ab948af1be372ac39da1a5109a3f5075a86c1b427d

Observation 1106ef4b-2dca-4fab-b580-21bb30ce0f6f · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.591430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:b8cc198e07429fc7d0d5880fb372bedd7740b265f07989db398631298b8cd9ed

Observation ab10d082-d293-410c-b093-1956c3539a4c · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Evalcrafter: Benchmarking and evaluating large video generation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.521794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:76c6fe8c15b6513ebccdcf99c4b7de0002434f16b1e6425942b6e4f11d3cdd6f

Observation da952127-2102-44db-975a-79b5132ffccd · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.525400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:8b11f0f4518d6f274ea7bf13da062625eae9ccaeb1b959978c04c9beb1c02267

Observation db2e0968-22bd-4b00-9db4-a9934965324e · outbound

This paper cites Learning transferable visual models from natural language supervision.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Learning transferable visual models from natural language supervision

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.515211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:921bb7febf3d5dba14eae1fd9917a8914e0a004cafede96bf9c205f7b34ec77b

Observation 7317b8e6-7035-4fe5-bc5a-421b8f8e8114 · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.653339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:4478651f182ec073c010e4c68437326d632bc516f839342267bc0d107a061862

Observation a66426f2-968e-4120-bc3a-b62884e0c948 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.670718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:b530b4f9d0f5310eee1d7246aacc5aab59bfe80b7ddda564a93b983e128a0287

Observation 3c80d119-73ab-4418-8651-03266d4a9214 · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Instructpix2pix: Learning to follow image editing instructions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.511061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:6fd547fa58fb80cc9931509e0804e8636ddf70d6333e5ba88362a07836bdc551

Observation 9fa9e7e8-9e24-4690-9c7c-343075a88130 · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.540123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:fb274f15ac8af75c8d92363f890ac8df933dcf6867f5ac7d90f418c9f7ac064a

Observation 6e5cccbd-516b-4e97-8fe9-1cc14bde0e01 · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:46.606255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:c6ea934eac6711a5aea5aff3f550f372443cfc8eb1a6be183db700afa31c0b20

Observation 8b80377b-c3f2-4c25-b3f5-aa0dab08222b · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Adding conditional control to text-to-image diffusion models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.523569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:aaa14fdc5866957b6784979533162c849ac9524d89e3f70d878a4cb033af75ae

Observation 56dda7ce-9b5f-414b-ad04-eb23813c1d3b · outbound

This paper cites T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.534972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:f51205f8ec285bf272fc3e81498422759c677d809324933e44de769c89952159

Observation c0a9aa0d-07a0-45b2-ad29-506a5f6e8acd · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Gligen: Open-set grounded text-to-image generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.476816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:b91ac57fb863fcc372265dd7827122ae0bf4633d9e377f29ff0c53882a90d5fb

Observation 0d733e11-866c-4d0d-8f41-c42ce17f3208 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.615070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:62d06e0a0f656cf7c1ae35708cf5d884a5f6f526ab2885f3017c439e5896e48e

Observation beb42f9b-670f-4520-8650-3bab9246490c · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.480037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:3f2fadc83599ceb4fa26ec0f817141cc0cfc83a5b775588b625503960bf90a87

Observation 441cbaee-5121-4a8a-8106-311eafb89e34 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.630251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:e696df568711ba5b668043f529149327879ecaaa48843a6d2ea7507fcc4a90a2

Observation ce856a71-1e92-473b-85c9-6468c9890ce8 · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.Advances in Neural Information Processing Systems, 36:7594–7611.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Videocomposer: Compositional video synthesis with motion controllability.Advances in Neural Information Processing Systems, 36:7594–7611

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.485977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:95ae4a5eee6a160fab367446d74fa8b663d77425eba1f167420d71fa642e1ae6

Observation d138ad65-c1a6-47cd-b2f7-ef89e511c0a8 · outbound

This paper cites Omnigen: Unified image generation.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Omnigen: Unified image generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.472365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:66bc7bad78c4e8fd5f5e686c650c92d0c1053b1c6d0689c6372f7d8be1bb0104

Observation e9c6ffa5-2c1e-4d62-8035-7e5b067cd766 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.583707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:c6eee09b733830cc1cd6431fc744092849ca66ca1d54efe720b3e6cf69b6ecfb

Observation e1414055-4be3-47e2-a770-7e9ccc0d408e · outbound

This paper cites Vace: All-in- one video creation and editing.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Vace: All-in- one video creation and editing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.480976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:20e761ed9beac623ee95e97920225f700f20b7708e798b52f9e459adccd535a6

Observation 8485f045-4432-4f0e-bf4f-46231a86e6a2 · outbound

This paper cites Character-aware models improve visual text rendering.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Character-aware models improve visual text rendering

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.482849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:42d3b544b42720fc879c772537ff63f9da2fd81d1b9d01b1a1c826a640d9b6d8

Observation c0a3d646-a242-4507-9269-5e9824a1fefe · outbound

This paper cites Glyphcontrol: Glyph conditional control for visual text generation.Advances in Neural Information Processing Systems, 36:44050–44066.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Glyphcontrol: Glyph conditional control for visual text generation.Advances in Neural Information Processing Systems, 36:44050–44066

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.465964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:2bca95a987705d9d74b4ca7dc9ab0aa5219c013a4b73b5db1d230b8507fcd851

Observation 487b97b5-e6d1-4f80-8d2c-0c55a89cfa75 · outbound

This paper cites AnyText: Multilingual Visual Text Generation And Editing.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm AnyText: Multilingual Visual Text Generation And Editing

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:46.586536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:22f428894c045dd6b16198c760aba937d4ede9a1d0e573e99342ea73919bb651

Observation 54256d8a-cafc-405b-bf4c-02fdbcffc5ec · outbound

This paper cites Textdiffuser: Diffusion models as text painters.Advances in Neural Information Processing Systems, 36: 9353–9387.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Textdiffuser: Diffusion models as text painters.Advances in Neural Information Processing Systems, 36: 9353–9387

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.467911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:221e62c4a68378eb211f358d0cee39c012a0646f855ace766a23b3d7514987df

Observation d8a22524-cabd-4247-a72a-6d023d450c46 · outbound

This paper cites Improving image generation with better captions.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Improving image generation with better captions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.484949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:1671f9df5b7f656d36ec7483be30fbf2039852feefe6b9445627c3b705fad30c

Observation 6bc2af46-c394-469b-8513-9d76ce622299 · outbound

This paper cites Visual prompting via image inpainting.Advances in neural information processing systems, 35: 25005–25017.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Visual prompting via image inpainting.Advances in neural information processing systems, 35: 25005–25017

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.515445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:5a7183304d2a10cf7feae13eda57fda0468ea6b5481d380cbd9a3c4e2b6db7f8

Observation 93cdd04e-ea6d-48a5-8a71-65b1ecd6a392 · outbound

This paper cites In-context learning unlocked for diffusion models.Advances in Neural Information Processing Systems, 36:8542–8562.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm In-context learning unlocked for diffusion models.Advances in Neural Information Processing Systems, 36:8542–8562

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.518281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:007936be9670a3639b0279bf3eb4b46d8629cff58d5e0962607e71abe85facab

Observation 96d56013-3980-4fe6-b534-6e54e29ee87b · outbound

This paper cites Context diffusion: In-context aware image generation.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Context diffusion: In-context aware image generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.505774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:07c0a35eff91a691fe02db0c3c62d3c1c33e49446b6712c67d3cb985694f34e6

Observation 7037d149-7862-4e18-b52c-1789bd1d74a0 · outbound

This paper cites Stable diffusion models are secretly good at visual in-context learning.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Stable diffusion models are secretly good at visual in-context learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.507841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:1b6ed3fe846bf15fa9cc217f72e774841154a39f71b11edd6d0f9dbab4cdef99

Observation 8639671c-cbe6-449e-84ec-23716b5e730f · outbound

This paper cites Visualcloze: A universal image generation framework via visual in-context learning.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Visualcloze: A universal image generation framework via visual in-context learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.459214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:6feb9f121ff0bdac917d1bc81904dfc2cc86b30be817ba4240e58241eeb04432

Observation 6f1e17da-0ba9-4c1e-aef2-05d6eeffa5e4 · outbound

This paper cites Realgeneral: Unifying visual generation via temporal in-context learning with video models.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Realgeneral: Unifying visual generation via temporal in-context learning with video models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.461133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:c1779e8c7277a4b7035636d5653b1d22f3b074c2c36a19feb98438f9dbcf666a

Observation bc3ed293-0cb5-48ed-9497-a0172144441f · outbound

This paper cites UNIC: Unified In-Context Video Editing.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm UNIC: Unified In-Context Video Editing

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:46.624365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:e5b68ccfc46201833ed852ed10204d7d373041cfcd9c2786d80cceb646b739fe

Observation e33e2eea-d06b-4c15-ba1d-632fae2e7cf2 · outbound

This paper cites Emu3.5: Native Multimodal Models are World Learners.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Emu3.5: Native Multimodal Models are World Learners

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.685305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:6f0fc376720fb2a899156bc15f724de6002ea8c1e6dd1ef8628c90122523e1de

Observation 3f9372ae-1500-49c6-ad5e-e7d2938c04e6 · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Show-o2: Improved Native Unified Multimodal Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.656154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:ee3fac4bd7c8c285c086c35b4d464fc08bb719e1bef0380261e30dd175e4a433

Observation f85f81c0-aaea-41f1-9d38-5b08f5451eb8 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Emerging Properties in Unified Multimodal Pretraining

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.665974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:1df002c0e8effb199bbd1e0fa35cc0bb2d4351f795c45fc770f8199494e3ea14

Observation f919d687-9f33-4bbc-91b7-f97d15e76730 · outbound

This paper cites Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.658675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:465c7291b52ca484e81e4506220bb3693b0a02531d42d4a453127a250c029e37

Observation 674feb39-9b0f-4fad-ab02-bfa10190aa6f · outbound

This paper cites DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:46.688203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:5881450889b99042d1d2a12822025491269d8815304c732656f8ce0818c68b0d

Observation f460fd29-d601-4093-8370-1dcdf5c1ee66 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.567810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:b4d313134a388960c62e74a2d182b08cf9d48b8dccbb2312bd5dcef54e600a1a

Observation 50ac9039-5608-455a-9201-3a517fe4f566 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.580939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:c65c5eafd3a4cfd1daf2aef4a5662ef91b6dcc95291c8208aad9d8dd0ea55801

Observation 203a4fda-74fe-4a2e-8b72-50a298f3f88a · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Emu3: Next-Token Prediction is All You Need

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.556714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:04665fb97c8416d68d77ebad2267d7a92d13a3d2681528d91020adcb23210e81

Observation 976438e6-5f3e-4571-b215-73b27d26f3fc · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.562530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:4bf75c0961cd8e3e1e9e062f51627aa69ab1dccc6851bed5075e00b295ab0767

Observation 973ff730-7b0c-42e0-950f-1fefdc5ec241 · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.463079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:2ef610d521cd016f6d9f67c867743f8fa7b71f415a705dbdba6ce1bc37d6ebb5

Observation 635ecd46-57b7-485b-815e-58226d205a13 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.633050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:0800eeb7751b6e41cfbcf7114304c910ab33f536e194a8cbbd900157df61900b

Observation c0ded2af-0e29-4fe6-94ac-7a7a5fe1a3b3 · outbound

This paper cites Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.457188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:bd052367ba3076eca5266804b0dfe517531ce2b6e53a700de2f79656da1f45a4

Observation 3b5808f0-1db0-4753-b66e-0bd699d46489 · outbound

This paper cites Lumina-image 2.0: A unified and efficient image generative framework.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Lumina-image 2.0: A unified and efficient image generative framework

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.517832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:1f05bd7ed22634fcb4476edabb9da24dee31bf20b21dec8ba9076921d4179eec

Observation dcf13539-75b9-4424-9a25-900f1f510917 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.578806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:c040fbc8a66a2f3bb00a37fff0e7869a339c8c0bf350cbea060654513a5b52ed

Observation 73fead70-c284-4fb2-b0e4-e7b2e61a92b8 · outbound

This paper cites HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.586663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:726d0f2a0a9c90b3a2346b8ad22e191aaf2f655b2ce2ea3c7d1a73f1454a98f8

Observation fcf0f092-8bf2-4fda-84ed-c51eeabdbd38 · outbound

This paper cites https://platform.openai.com/docs/ models/gpt-image-1.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm https://platform.openai.com/docs/ models/gpt-image-1

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.519833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:9dbda219e94e3f43d96c84f6eb7211d7070fc503a7c649d7485fb4be5da49c38

Observation 40ef2f1a-2b6e-4061-bc22-cacf32f600b5 · outbound

This paper cites Seedream 3.0 Technical Report.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm Seedream 3.0 Technical Report

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.668339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:af7fc1509ce31d0be5a3cb0f3c469e0de8f98e781c687b37e4965f5f0978217e

Observation a588c5d5-47df-4ba9-ab1f-f46ff7366eba · outbound

This paper cites https://platform.openai.com/docs/ models/gpt-image-2.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm https://platform.openai.com/docs/ models/gpt-image-2

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.453633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:e9758f9917aa4626eb978c1925e14ac6c931325f9405da0dc0df9355c56290eb

Observation ca13dcc8-3372-4dde-a174-006b76214fb5 · outbound

This paper cites https://seed.bytedance.com/en/seedream5_ 0_lite.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm https://seed.bytedance.com/en/seedream5_ 0_lite

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T14:13:49.455306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:fa1e2bed68495cfb3b13620b579d1be6e9ff8d5ed5c35a44b2e84dfe464c6470

Observation d3ba91d1-bdc0-461f-965b-3d00e4a002fd · outbound

This paper cites minor imperfections only.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm minor imperfections only

Reference 66

Resolution
malformed identifier
arxiv_id, observed 2026-07-01T14:05:46.591957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:80d5f228062918c4f680cd6a3ae0c607f7bb5c96f395acdf32d873d0f2d260b4

Pith citing papers

No inbound Pith citation observations are available.