Pith. sign in

Paper Citation Record · LEDGER

Learning Visual Generative Priors without Text

As of 21 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 1 inbound Pith citation observation for arXiv:2412.07767.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.07767 v3

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:35:53.182869Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:54.032525Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:48:57.600976Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1d945ba8-876b-4e0e-8b72-7e1883ae83a4 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Learning Visual Generative Priors without Text Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.351574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.850914Z digest=sha256:1a1ad20a3c27cef83cce339c43cada247f6d6eef3e08d9c9bf42c04eec4549bd

Observation 25716040-1922-4de3-a98b-34a5dbf5ced6 · outbound

This paper cites Improving image generation with better captions.

Learning Visual Generative Priors without Text Improving image generation with better captions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.334199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.856875Z digest=sha256:acad153e65050cda354d56308bd0ecfb9910b925896867ca2d0abee9f4e9ce9d

Observation 215372b7-5391-46a4-858a-cd5616c648a2 · outbound

This paper cites an unresolved cited work.

Learning Visual Generative Priors without Text Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:35:54.314943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.862570Z digest=sha256:e68a2893af2615db0f85894c974959bf7f9ff97da87a4a647da7575c8eef6bdb

Observation 5c627f0a-2e28-4fb5-9a46-f208b9997a9b · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Learning Visual Generative Priors without Text Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:52.868174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:52.868174Z digest=sha256:fc9db30a5f02cac33a0571675cfe027f4312338ba0c106e994aa993359544ea6

Observation 030d34f8-0757-43a1-adc3-e0f4ef02a488 · outbound

This paper cites Coyo- 700m: Image-text pair dataset.

Learning Visual Generative Priors without Text Coyo- 700m: Image-text pair dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.297671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.874335Z digest=sha256:56086514b57e07909154902b6efbcc9df0200ab0cf269bb4ae55d0a3b4ad6dac

Observation 4d3dc9be-d68c-4ea8-b026-6d1578225ea8 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Learning Visual Generative Priors without Text Emerg- ing properties in self-supervised vision transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.278752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.879831Z digest=sha256:6f10fc49cb7fdb5fd2a7dda61f3ae709ac352d6d121d9f048a248cd1fa9472c4

Observation c68f9622-f5d6-47ec-9728-8c632b8de87f · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Learning Visual Generative Priors without Text PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:52.886220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:52.886220Z digest=sha256:f90f2b4f35a879738e1682f245782e38df8a255135e87094ccc2d164b7b3ede4

Observation 0ca1cbc6-0452-424b-b0d4-f7e6b680ebda · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

Learning Visual Generative Priors without Text PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:52.891840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:52.891840Z digest=sha256:122d3a8b42892238b99c02e8126d7c36da6fb462add9e2bdbfc51a58bff21c42

Observation 8b318a45-916f-4718-bb10-e60942ce991e · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Learning Visual Generative Priors without Text A simple framework for contrastive learning of visual representations

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.259394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.897571Z digest=sha256:c12be0abff5780725852693f2ff7dbe6504618c6d7e1d1fd9021d139e50af8ce

Observation b6aed268-7560-421a-90b6-b24e23faf3db · outbound

This paper cites Improved Baselines with Momentum Contrastive Learning.

Learning Visual Generative Priors without Text Improved Baselines with Momentum Contrastive Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:52.903113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:52.903113Z digest=sha256:64b07ec5f4c6545e70830e4722f8f5d022df89fc10f71f59fd8eef217cac8f00

Observation 23a1bd5f-b2ac-43c2-adca-44ca3f591dab · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Learning Visual Generative Priors without Text Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.239267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.909129Z digest=sha256:9fc34fc57774e0dce565dd5489466dbc72644a960323062f4d1a789d35231c02

Observation 8ccd94ba-378d-4313-912f-7d52620b2d0f · outbound

This paper cites Objaverse-xl: A universe of 10m+ 3d objects.

Learning Visual Generative Priors without Text Objaverse-xl: A universe of 10m+ 3d objects

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.219875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.914329Z digest=sha256:ed79e73d502110e45542564b381914cca691630be070f2cf71f3c70ecc06b192

Observation 40ee1db0-f12b-4b93-8583-780d85bd2e0d · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

Learning Visual Generative Priors without Text Objaverse: A universe of annotated 3d objects

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.199838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.919427Z digest=sha256:20266d6d3795dc718734b5d30d90c6482e6708f33f20f1ab929f354a35f3ed8a

Observation 0b6d823e-9fe8-46ed-9d14-5c9720384945 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Learning Visual Generative Priors without Text Imagenet: A large-scale hierarchical image database

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.181995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.924691Z digest=sha256:1ffbe8c86bc9d6bea5bc71d4590ef20cbea0885b37c850454c7cd46ce821cbfb

Observation f2f0e166-a948-4f40-84cf-42195e641b2a · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Learning Visual Generative Priors without Text BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:52.929930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:52.929930Z digest=sha256:67d67add620f7aca18c944e91a00f841f0e7b7f76bf45bb13f34fa685a541c39

Observation a952aaa3-8359-4043-85a6-f34fcd884889 · outbound

This paper cites Google scanned objects: A high- quality dataset of 3d scanned household items.

Learning Visual Generative Priors without Text Google scanned objects: A high- quality dataset of 3d scanned household items

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.163910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.935442Z digest=sha256:385f85210437e21c4bdef7e240b393c6d0e1cc4806391e2f2e293d5c335390b9

Observation e27b06b4-2db8-4ce4-872e-e06774b1663a · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Learning Visual Generative Priors without Text Scaling rectified flow transformers for high-resolution image synthesis

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.145730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.941166Z digest=sha256:5f40e7b53e1aa3b5fd5b74fa4ec4c53316f55ea96991e94de0d31ffbc013de42

Observation 7e7ad258-89e8-4cb5-bc26-8d427654350e · outbound

This paper cites Geneval: An object-focused framework for evaluating text- to-image alignment.

Learning Visual Generative Priors without Text Geneval: An object-focused framework for evaluating text- to-image alignment

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.124302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.946957Z digest=sha256:1d24eed66a84ce2e319e1c3a1d948cbd7e05ff71cd3cf1bde73a287769037215

Observation 7c8d2470-d1af-4f34-963c-f2118b7014d7 · outbound

This paper cites Generative adversarial nets.

Learning Visual Generative Priors without Text Generative adversarial nets

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.102115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.953207Z digest=sha256:c1cab3434bf213743f8defd5756e2bc98ebf251db0d99e886643e028d70726b1

Observation edfb1d48-e66e-47db-b5b5-524c2554be02 · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

Learning Visual Generative Priors without Text Momentum contrast for unsupervised visual representation learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.082657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.958271Z digest=sha256:568833c7e0f72f5c38cbcb0d6ffe43e6a56728a1f49262b8f4a06674ef45157a

Observation 0836d6c1-af43-4844-ae46-d0446207cbfa · outbound

This paper cites Masked autoencoders are scalable vision learners.

Learning Visual Generative Priors without Text Masked autoencoders are scalable vision learners

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.063766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.963887Z digest=sha256:1c269d0eaa954838bd92771df68eb7fb579a9379d7bd6bbc0fc1b02f356e9cfc

Observation 405b64bf-e0a5-42f3-8161-d0339e706285 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Learning Visual Generative Priors without Text Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.041970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.970785Z digest=sha256:2a56428151de7df24b560617db13e7e6053fba0b4fea80c23a1a0f51d278fc9b

Observation 1e7432ab-46af-4820-aad7-d656040ae011 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Learning Visual Generative Priors without Text ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:52.976674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:52.976674Z digest=sha256:fb74564a5500cf790d0023062a8d6c66b2c96f276d58b97fd2ad62f5abf51e25

Observation d6e0c049-b25c-41d2-a972-85d6e4f68537 · outbound

This paper cites Scaling up gans for text-to-image synthesis.

Learning Visual Generative Priors without Text Scaling up gans for text-to-image synthesis

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:54.019045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.982558Z digest=sha256:1c2b1336ffc2fc4b68dd6a442d930d41fa864f0b5112430ac301f2b413af9553

Observation 41d1822f-0d7a-4dc0-b31e-db0b59bc7fbc · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.

Learning Visual Generative Priors without Text 3d gaussian splatting for real-time radiance field rendering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:52.987975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:52.987975Z digest=sha256:94b2e321cb32edfeb124ba35d0c9f9ca5adf84a4cb4d0c119a7fe685a5f3edd5

Observation 252d18a7-0450-4c81-84c5-d7224112b333 · outbound

This paper cites Auto-encoding variational bayes.

Learning Visual Generative Priors without Text Auto-encoding variational bayes

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.980832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.994000Z digest=sha256:ce1b7c768a5e18c26c68e3064d55024c88d0bb3126d09a86616eb55e135e586d

Observation 39d1a505-5cce-47dd-a320-450f591d5e66 · outbound

This paper cites Segment anything.

Learning Visual Generative Priors without Text Segment anything

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.959523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:52.999924Z digest=sha256:19187856f40c5f279335563995b9434a17f80305c3131c770249be36aac7aace

Observation 0eb16328-c268-403d-b82c-da90bf0cbc6f · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

Learning Visual Generative Priors without Text Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:53.005761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:53.005761Z digest=sha256:a93931197bfb5c3b9dc7c48e08cb6b74c34ecf044fcd5642c50b65c9f5e34f68

Observation 5083606b-b53a-4e2f-9945-1ad18d0f7e3c · outbound

This paper cites Return of Unconditional Generation: A Self-supervised Representation Generation Method.

Learning Visual Generative Priors without Text Return of Unconditional Generation: A Self-supervised Representation Generation Method

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:53.011786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:53.011786Z digest=sha256:e48ca96c35c283b28f971c05d7a4632435df50fa0aaa4e55dd2ca36029de0e4a

Observation f3eb7c07-0b98-441d-8f8e-3b591ef34c24 · outbound

This paper cites Microsoft coco: Common objects in context.

Learning Visual Generative Priors without Text Microsoft coco: Common objects in context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:53.017170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:53.017170Z digest=sha256:099bc8407ca6557b7920cfca0fba205f4a33c43905ff1e824b02c63a75892202

Observation 57614c2b-4f73-4fc6-90b9-3005238bab52 · outbound

This paper cites Zero-1-to-3: Zero-shot one image to 3d object.

Learning Visual Generative Priors without Text Zero-1-to-3: Zero-shot one image to 3d object

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.927243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.022458Z digest=sha256:accda11015c38e03a8f458fe0b99e3076afe0d8079fc40e8ff66e5dd875de4d0

Observation 352dccec-9e8f-4fef-a273-a014758f3a54 · outbound

This paper cites SyncDreamer: Generating Multiview-consistent Images from a Single-view Image.

Learning Visual Generative Priors without Text SyncDreamer: Generating Multiview-consistent Images from a Single-view Image

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:53.027738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:53.027738Z digest=sha256:610475d9787f92733a36a95b68476eaaa6ef5223b52c44945c69c86e83cd6dd1

Observation b91ec6ee-f9e1-4c33-8374-d9189145d155 · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

Learning Visual Generative Priors without Text OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:53.033311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:53.033311Z digest=sha256:ad6bac70b211d0697b98249285991a0b1c86f2c78d4c53d267c98cebb3ce0fff

Observation f1b1eaaa-6d4d-48cf-8755-e0af5b2bab42 · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

Learning Visual Generative Priors without Text GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:53.038724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:53.038724Z digest=sha256:96f13d97cb616935e76997224f70d3ca925fd5bb9ba784f430cb71e40a3df01c

Observation 1107465f-4a6a-495c-9773-a68b8189b8cd · outbound

This paper cites Jour- neydb: A benchmark for generative image understanding.

Learning Visual Generative Priors without Text Jour- neydb: A benchmark for generative image understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.908809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.044600Z digest=sha256:6e8e83147bfcbe2f764608883c55679ecc7d665eb9d8d9786074f0bdf83fc25b

Observation 247b4d85-8c32-41bd-8f03-9cab8933ec0a · outbound

This paper cites Scalable di ffusion models with transformers.

Learning Visual Generative Priors without Text Scalable di ffusion models with transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.890345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.050468Z digest=sha256:dd709654c2c73f9ee9f89c35e5af047ecab5d2d68352199e236f0bfaa47e583d

Observation d8eea6a9-ad9c-4a23-8de3-28d616b35ca9 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Learning Visual Generative Priors without Text SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:53.055569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:53.055569Z digest=sha256:66481a76e6ec662c03e174564e92b9aba2b39584f905fd6f1ead386eeb253fb4

Observation cb932347-0583-435a-aff5-2d6b93c2700f · outbound

This paper cites Improving language understanding by gener- ative pre-training.

Learning Visual Generative Priors without Text Improving language understanding by gener- ative pre-training

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.871507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.061126Z digest=sha256:f69a3db6a8c5fcfd84e69bd3a65db607e01274effaf1367baac0f8588558b9c2

Observation 49299552-d16a-4283-9653-3da96d049503 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Learning Visual Generative Priors without Text Learning transferable visual models from natural language supervi- sion

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:53.066528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:53.066528Z digest=sha256:ee315f1f493cfa5e81fcddc00e8ca86f3fdac749d434d3d547343338010b94ba

Observation 8956747d-34bf-431b-ab90-c66ae0e30d41 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Learning Visual Generative Priors without Text Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.840092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.071869Z digest=sha256:0f0bab43363bba0fe50f1432904ab6a85a338e99f59f9945669ea5e2d6d2926a

Observation c47492fc-7c79-42fe-89fe-9337a2f8e67b · outbound

This paper cites Zero-shot text-to-image generation.

Learning Visual Generative Priors without Text Zero-shot text-to-image generation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.820281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.077036Z digest=sha256:a8d8697929f24d0df9eaa9638e250db8a45ad42a61bc9adb64d1c979b2583594

Observation 9db0274d-b3d4-4501-aaed-878f0edba2c5 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Learning Visual Generative Priors without Text Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:53.082291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:53.082291Z digest=sha256:c8b46eda9376c2a19de7032dbd992b8085ad485be0a323d8c9c3db68707a0ef4

Observation b08c0b49-8798-4296-ae8b-9235a6863712 · outbound

This paper cites Variational inference with normalizing flows.

Learning Visual Generative Priors without Text Variational inference with normalizing flows

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.799506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.087752Z digest=sha256:5dcfe62a3b389e514a8454644096c36106b419b7515ea6d7dc648eace189622b

Observation fa2b0b9f-6746-4999-a515-fba783228f16 · outbound

This paper cites High-resolution image synthesis with latent di ffusion models.

Learning Visual Generative Priors without Text High-resolution image synthesis with latent di ffusion models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.779551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.093238Z digest=sha256:66155220a362991dba3be8fef2d86fd6e13e5df2efdca96334ba81c278ba9fe2

Observation 7a026372-91dd-4613-ba38-d5731b24ff8f · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Learning Visual Generative Priors without Text Photorealistic text-to-image diffusion models with deep language understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.759498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.098657Z digest=sha256:8930563a7b7727959b75ba70f980ec57025b34cfa280e2f2ba6a6f5213c14e98

Observation 3aac7f8f-6570-467f-ab04-76944db24502 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Learning Visual Generative Priors without Text Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.741179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.103989Z digest=sha256:df3d176f662882aaa8af0fa1d51e942acc47e8588625b2dea006bc69890665cb

Observation b0dbe713-91e2-41cc-8520-43c6ec7ba584 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Learning Visual Generative Priors without Text UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:53.109361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:53.109361Z digest=sha256:15408609a936056f859d3e9036f9afcb88ee275309285d2e60ca48a2dfffbf9f

Observation 2c57ad8b-30ed-4671-bfe4-1c28995b644c · outbound

This paper cites Lgm: Large multi-view gaussian model for high-resolution 3d content creation.

Learning Visual Generative Priors without Text Lgm: Large multi-view gaussian model for high-resolution 3d content creation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.722725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.115720Z digest=sha256:af4b671e0274e3b58b4297822023954bea6b91453f48a51667826d582351c829

Observation cbec7c4c-8f1d-48c5-a12a-f742ef098caa · outbound

This paper cites Kolors: E ffective training of diffusion model for photorealistic text-to-image synthesis.

Learning Visual Generative Priors without Text Kolors: E ffective training of diffusion model for photorealistic text-to-image synthesis

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.701342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.121162Z digest=sha256:5a0f380ed33771a8ce671e85bec1bc031280cd9c0d70f9711f2140d438bae118

Observation 69832b45-fdf7-4f7c-a319-e092ccadf305 · outbound

This paper cites Fvd: A new metric for video generation.

Learning Visual Generative Priors without Text Fvd: A new metric for video generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:53.126590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:53.126590Z digest=sha256:209157df57a49bad91e0422a940ab0bc278f45b75611bea628eb16834291b84d

Observation 9f8c7180-f9ce-444f-b7f6-411c4e4e1db9 · outbound

This paper cites Attention is all you need.

Learning Visual Generative Priors without Text Attention is all you need

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.669732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.132056Z digest=sha256:d5555f6663d98b9d3111435d9f9c49b38f97fc674b911a82f2c5820b51ca9be3

Observation 20d411e1-061e-4c66-91b4-6860e45bcab1 · outbound

This paper cites Qihoo-T2X: An Efficient Proxy-Tokenized Diffusion Transformer for Text-to-Any-Task.

Learning Visual Generative Priors without Text Qihoo-T2X: An Efficient Proxy-Tokenized Diffusion Transformer for Text-to-Any-Task

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:53.137350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:53.137350Z digest=sha256:389122ff272c446b239c4c98ba9cb915668fa4093f754e94f3b8cf8901f63b50

Observation 4ec2e57a-9826-4e5e-9b22-a6baa4286fb9 · outbound

This paper cites SimMIM: A simple framework for masked image modeling.

Learning Visual Generative Priors without Text SimMIM: A simple framework for masked image modeling

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.650647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.142881Z digest=sha256:1ef80600ac3291edc96f8d186d1655313f60f993d97a0f608b88073332ff7b4c

Observation 8b9260d8-eb21-4ee8-8b02-03dca695e221 · outbound

This paper cites Dynamicrafter: Animating 10 open-domain images with video di ffusion priors.

Learning Visual Generative Priors without Text Dynamicrafter: Animating 10 open-domain images with video di ffusion priors

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.631143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.148676Z digest=sha256:7e1fa3ca0144046f05cec2e6aab6a051d217d9a27124a91bfcb69f22cea33506

Observation 8958e374-c7b9-4df4-845e-224fe08e0906 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Learning Visual Generative Priors without Text Msr-vtt: A large video description dataset for bridging video and language

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.613145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.154428Z digest=sha256:7d3d1c8c99bd2fca7c7f2af6d5019254b4c4fce5c7fcf284f7c4dc1715210816

Observation b281cc8d-d81a-44d6-b615-b04c596ee4a8 · outbound

This paper cites GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation.

Learning Visual Generative Priors without Text GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:53.159771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:53.159771Z digest=sha256:2b7a1ec5e354745a013823c94a78d555e4a034a6fc661422a78fdd66679ee058

Observation 27c22cad-938e-4401-a564-2f971e68ce15 · outbound

This paper cites Raphael: Text-to- image generation via large mixture of di ffusion paths.

Learning Visual Generative Priors without Text Raphael: Text-to- image generation via large mixture of di ffusion paths

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.593802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.165280Z digest=sha256:768d23b1561cc49309f492e27158d7481df0657067ae022ff131d9a01d1d3f8f

Observation 48b7d5de-a9aa-4b4f-b56d-bf8a3cf3b119 · outbound

This paper cites Open-sora: Democratizing e fficient video production for all.

Learning Visual Generative Priors without Text Open-sora: Democratizing e fficient video production for all

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.574093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.171014Z digest=sha256:e1230ee1fdc753106f560b05f45443b04c4f41933ec894132e806ab1e3f6299f

Observation 5e1b01ed-c46a-4fa0-8795-1961dcde9557 · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

Learning Visual Generative Priors without Text Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:53.176731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:53.176731Z digest=sha256:134fa8fe147e81c754d5f184e82ae0d5b49a9385a487da63535ea7b390b48eec

Observation a0227ccf-21b2-41c3-a85b-daf9378fd498 · outbound

This paper cites golden sunset shines on the top of snow-capped mountains, with small villages at its foot and surrounding buildings.

Learning Visual Generative Priors without Text golden sunset shines on the top of snow-capped mountains, with small villages at its foot and surrounding buildings

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:35:53.554892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:35:53.182869Z digest=sha256:b278603e50ab9a4e9c8d5caedeaf9e1605178f13575111aee65e3f35a54fdcfb

Pith citing papers

Observation 05868bab-5a3c-4920-9e7d-12d77b8175ba · inbound

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation cites this paper.

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation Learning Visual Generative Priors without Text

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:48:57.707197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T12:48:54.032525Z digest=sha256:de712fc30abb3214feaa8162328a119e418d3bf45c8c857d3870e1fd8ba09ff5