Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T18:35:53.182869Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 1 inbound Pith citation observation for arXiv:2412.07767.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T18:35:53.182869Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:54.032525Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T12:48:57.600976Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1d945ba8-876b-4e0e-8b72-7e1883ae83a4 · outbound
Learning Visual Generative Priors without Text Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 25716040-1922-4de3-a98b-34a5dbf5ced6 · outbound
Learning Visual Generative Priors without Text Improving image generation with better captions
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 215372b7-5391-46a4-858a-cd5616c648a2 · outbound
Learning Visual Generative Priors without Text Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5c627f0a-2e28-4fb5-9a46-f208b9997a9b · outbound
Learning Visual Generative Priors without Text Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 030d34f8-0757-43a1-adc3-e0f4ef02a488 · outbound
Learning Visual Generative Priors without Text Coyo- 700m: Image-text pair dataset
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4d3dc9be-d68c-4ea8-b026-6d1578225ea8 · outbound
Learning Visual Generative Priors without Text Emerg- ing properties in self-supervised vision transformers
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c68f9622-f5d6-47ec-9728-8c632b8de87f · outbound
Learning Visual Generative Priors without Text PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ca1cbc6-0452-424b-b0d4-f7e6b680ebda · outbound
Learning Visual Generative Priors without Text PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b318a45-916f-4718-bb10-e60942ce991e · outbound
Learning Visual Generative Priors without Text A simple framework for contrastive learning of visual representations
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b6aed268-7560-421a-90b6-b24e23faf3db · outbound
Learning Visual Generative Priors without Text Improved Baselines with Momentum Contrastive Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23a1bd5f-b2ac-43c2-adca-44ca3f591dab · outbound
Learning Visual Generative Priors without Text Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8ccd94ba-378d-4313-912f-7d52620b2d0f · outbound
Learning Visual Generative Priors without Text Objaverse-xl: A universe of 10m+ 3d objects
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 40ee1db0-f12b-4b93-8583-780d85bd2e0d · outbound
Learning Visual Generative Priors without Text Objaverse: A universe of annotated 3d objects
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0b6d823e-9fe8-46ed-9d14-5c9720384945 · outbound
Learning Visual Generative Priors without Text Imagenet: A large-scale hierarchical image database
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f2f0e166-a948-4f40-84cf-42195e641b2a · outbound
Learning Visual Generative Priors without Text BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a952aaa3-8359-4043-85a6-f34fcd884889 · outbound
Learning Visual Generative Priors without Text Google scanned objects: A high- quality dataset of 3d scanned household items
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e27b06b4-2db8-4ce4-872e-e06774b1663a · outbound
Learning Visual Generative Priors without Text Scaling rectified flow transformers for high-resolution image synthesis
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7e7ad258-89e8-4cb5-bc26-8d427654350e · outbound
Learning Visual Generative Priors without Text Geneval: An object-focused framework for evaluating text- to-image alignment
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7c8d2470-d1af-4f34-963c-f2118b7014d7 · outbound
Learning Visual Generative Priors without Text Generative adversarial nets
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation edfb1d48-e66e-47db-b5b5-524c2554be02 · outbound
Learning Visual Generative Priors without Text Momentum contrast for unsupervised visual representation learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0836d6c1-af43-4844-ae46-d0446207cbfa · outbound
Learning Visual Generative Priors without Text Masked autoencoders are scalable vision learners
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 405b64bf-e0a5-42f3-8161-d0339e706285 · outbound
Learning Visual Generative Priors without Text Gans trained by a two time-scale update rule converge to a local nash equilibrium
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1e7432ab-46af-4820-aad7-d656040ae011 · outbound
Learning Visual Generative Priors without Text ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6e0c049-b25c-41d2-a972-85d6e4f68537 · outbound
Learning Visual Generative Priors without Text Scaling up gans for text-to-image synthesis
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 41d1822f-0d7a-4dc0-b31e-db0b59bc7fbc · outbound
Learning Visual Generative Priors without Text 3d gaussian splatting for real-time radiance field rendering
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 252d18a7-0450-4c81-84c5-d7224112b333 · outbound
Learning Visual Generative Priors without Text Auto-encoding variational bayes
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 39d1a505-5cce-47dd-a320-450f591d5e66 · outbound
Learning Visual Generative Priors without Text Segment anything
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0eb16328-c268-403d-b82c-da90bf0cbc6f · outbound
Learning Visual Generative Priors without Text Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5083606b-b53a-4e2f-9945-1ad18d0f7e3c · outbound
Learning Visual Generative Priors without Text Return of Unconditional Generation: A Self-supervised Representation Generation Method
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3eb7c07-0b98-441d-8f8e-3b591ef34c24 · outbound
Learning Visual Generative Priors without Text Microsoft coco: Common objects in context
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57614c2b-4f73-4fc6-90b9-3005238bab52 · outbound
Learning Visual Generative Priors without Text Zero-1-to-3: Zero-shot one image to 3d object
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 352dccec-9e8f-4fef-a273-a014758f3a54 · outbound
Learning Visual Generative Priors without Text SyncDreamer: Generating Multiview-consistent Images from a Single-view Image
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b91ec6ee-f9e1-4c33-8374-d9189145d155 · outbound
Learning Visual Generative Priors without Text OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1b1eaaa-6d4d-48cf-8755-e0af5b2bab42 · outbound
Learning Visual Generative Priors without Text GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1107465f-4a6a-495c-9773-a68b8189b8cd · outbound
Learning Visual Generative Priors without Text Jour- neydb: A benchmark for generative image understanding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 247b4d85-8c32-41bd-8f03-9cab8933ec0a · outbound
Learning Visual Generative Priors without Text Scalable di ffusion models with transformers
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d8eea6a9-ad9c-4a23-8de3-28d616b35ca9 · outbound
Learning Visual Generative Priors without Text SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb932347-0583-435a-aff5-2d6b93c2700f · outbound
Learning Visual Generative Priors without Text Improving language understanding by gener- ative pre-training
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 49299552-d16a-4283-9653-3da96d049503 · outbound
Learning Visual Generative Priors without Text Learning transferable visual models from natural language supervi- sion
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8956747d-34bf-431b-ab90-c66ae0e30d41 · outbound
Learning Visual Generative Priors without Text Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c47492fc-7c79-42fe-89fe-9337a2f8e67b · outbound
Learning Visual Generative Priors without Text Zero-shot text-to-image generation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9db0274d-b3d4-4501-aaed-878f0edba2c5 · outbound
Learning Visual Generative Priors without Text Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b08c0b49-8798-4296-ae8b-9235a6863712 · outbound
Learning Visual Generative Priors without Text Variational inference with normalizing flows
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fa2b0b9f-6746-4999-a515-fba783228f16 · outbound
Learning Visual Generative Priors without Text High-resolution image synthesis with latent di ffusion models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7a026372-91dd-4613-ba38-d5731b24ff8f · outbound
Learning Visual Generative Priors without Text Photorealistic text-to-image diffusion models with deep language understanding
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3aac7f8f-6570-467f-ab04-76944db24502 · outbound
Learning Visual Generative Priors without Text Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b0dbe713-91e2-41cc-8520-43c6ec7ba584 · outbound
Learning Visual Generative Priors without Text UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c57ad8b-30ed-4671-bfe4-1c28995b644c · outbound
Learning Visual Generative Priors without Text Lgm: Large multi-view gaussian model for high-resolution 3d content creation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cbec7c4c-8f1d-48c5-a12a-f742ef098caa · outbound
Learning Visual Generative Priors without Text Kolors: E ffective training of diffusion model for photorealistic text-to-image synthesis
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 69832b45-fdf7-4f7c-a319-e092ccadf305 · outbound
Learning Visual Generative Priors without Text Fvd: A new metric for video generation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f8c7180-f9ce-444f-b7f6-411c4e4e1db9 · outbound
Learning Visual Generative Priors without Text Attention is all you need
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 20d411e1-061e-4c66-91b4-6860e45bcab1 · outbound
Learning Visual Generative Priors without Text Qihoo-T2X: An Efficient Proxy-Tokenized Diffusion Transformer for Text-to-Any-Task
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ec2e57a-9826-4e5e-9b22-a6baa4286fb9 · outbound
Learning Visual Generative Priors without Text SimMIM: A simple framework for masked image modeling
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8b9260d8-eb21-4ee8-8b02-03dca695e221 · outbound
Learning Visual Generative Priors without Text Dynamicrafter: Animating 10 open-domain images with video di ffusion priors
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8958e374-c7b9-4df4-845e-224fe08e0906 · outbound
Learning Visual Generative Priors without Text Msr-vtt: A large video description dataset for bridging video and language
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b281cc8d-d81a-44d6-b615-b04c596ee4a8 · outbound
Learning Visual Generative Priors without Text GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27c22cad-938e-4401-a564-2f971e68ce15 · outbound
Learning Visual Generative Priors without Text Raphael: Text-to- image generation via large mixture of di ffusion paths
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 48b7d5de-a9aa-4b4f-b56d-bf8a3cf3b119 · outbound
Learning Visual Generative Priors without Text Open-sora: Democratizing e fficient video production for all
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5e1b01ed-c46a-4fa0-8795-1961dcde9557 · outbound
Learning Visual Generative Priors without Text Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0227ccf-21b2-41c3-a85b-daf9378fd498 · outbound
Learning Visual Generative Priors without Text golden sunset shines on the top of snow-capped mountains, with small villages at its foot and surrounding buildings
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 05868bab-5a3c-4920-9e7d-12d77b8175ba · inbound
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation Learning Visual Generative Priors without Text
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.