Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:22:55.635316Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 1 inbound Pith citation observation for arXiv:2412.06774.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:22:55.635316Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:54:09.796839Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T18:54:12.830432Z
89 of 89 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fa53fbcf-cfba-451d-b919-ac685e9b21d2 · outbound
Visual Lexicon: Rich Image Features in Language Space Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df5898d1-f414-46de-844d-3616b039a90c · outbound
Visual Lexicon: Rich Image Features in Language Space BEit: BERT pre-training of image transformers
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8110e71e-2571-433c-812b-960dd4e6a763 · outbound
Visual Lexicon: Rich Image Features in Language Space Label-efficient se- mantic segmentation with diffusion models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11a12ab5-d0aa-4f21-b0e2-42b25e537971 · outbound
Visual Lexicon: Rich Image Features in Language Space Generalized denoising auto-encoders as generative models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bf1a14f-f047-4925-b72f-44ddb677ce87 · outbound
Visual Lexicon: Rich Image Features in Language Space Improving image generation with better captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 187cf068-12cc-4afb-89fd-f1b572370631 · outbound
Visual Lexicon: Rich Image Features in Language Space PaliGemma: A versatile 3B VLM for transfer
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7233933-ab07-44fb-91cf-bdef1edd0a01 · outbound
Visual Lexicon: Rich Image Features in Language Space Language Models are Few-Shot Learners
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34f3191c-b220-4d09-b037-40fae470963c · outbound
Visual Lexicon: Rich Image Features in Language Space Emerg- ing properties in self-supervised vision transformers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cf3a614-1b2c-4dd0-b362-f39cbf9c9fa8 · outbound
Visual Lexicon: Rich Image Features in Language Space Generative pre- training from pixels
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed4090ed-fc6a-48d9-acaf-358c37ae14e3 · outbound
Visual Lexicon: Rich Image Features in Language Space Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4471d54f-f548-4395-83ad-6a7ecc68b03c · outbound
Visual Lexicon: Rich Image Features in Language Space PaLI: A Jointly-Scaled Multilingual Language-Image Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ed146b8-028f-4577-a6e2-05d74d7da021 · outbound
Visual Lexicon: Rich Image Features in Language Space Deconstructing Denoising Diffusion Models for Self-Supervised Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7541b5ae-c73c-49bb-853f-026e0c757138 · outbound
Visual Lexicon: Rich Image Features in Language Space Reproducible scal- ing laws for contrastive language-image learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f2108ee-e53b-4aaa-9503-3a674420d1c9 · outbound
Visual Lexicon: Rich Image Features in Language Space Imagenet: A large-scale hierarchical image database
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03c98d72-164b-46e9-8f91-76a18d6c12ef · outbound
Visual Lexicon: Rich Image Features in Language Space Large scale adversarial representation learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfbf0931-9762-45ec-a0db-b6e1c13dcffa · outbound
Visual Lexicon: Rich Image Features in Language Space An image is worth 16x16 words: Transformers for image recognition at scale
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ccb5774-6814-46c8-833f-2be2bdda1bf9 · outbound
Visual Lexicon: Rich Image Features in Language Space A new algorithm for data compression
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation af91a1d3-e1e5-4ffb-918d-634b705575b9 · outbound
Visual Lexicon: Rich Image Features in Language Space An image is worth one word: Personalizing text-to-image gen- eration using textual inversion
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1d414d07-28c6-4f8d-89ae-da66703c8d72 · outbound
Visual Lexicon: Rich Image Features in Language Space Instructcv: Instruction- tuned text-to-image diffusion models as vision generalists
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9ebe5898-7155-478e-9f7c-984e83ac06bf · outbound
Visual Lexicon: Rich Image Features in Language Space Imagebind: One embedding space to bind them all
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fcdbec4-24b4-4a5d-9cd8-4672cdd60922 · outbound
Visual Lexicon: Rich Image Features in Language Space Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3d9bdb7a-b3b2-4aba-8c93-62428a98577e · outbound
Visual Lexicon: Rich Image Features in Language Space Vizwiz grand challenge: Answering visual questions from blind people
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfda47b2-9d57-4352-bc29-30acf26db66c · outbound
Visual Lexicon: Rich Image Features in Language Space Masked autoencoders are scalable vision learners
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7bc44f81-38ce-4805-bc07-4364fc897ec9 · outbound
Visual Lexicon: Rich Image Features in Language Space Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec23e6a0-1c94-4aab-a2a6-dc8159441e07 · outbound
Visual Lexicon: Rich Image Features in Language Space Autoencoders, mini- mum description length and helmholtz free energy.Advances in neural information processing systems, 6, 1993
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f8c2c37b-6659-46f5-b568-2aeafbbc2a9a · outbound
Visual Lexicon: Rich Image Features in Language Space Classifier-free diffusion guidance
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38770467-6fb6-47fe-a205-3d69acc679aa · outbound
Visual Lexicon: Rich Image Features in Language Space Denoising dif- fusion probabilistic models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd3f989b-bf23-49db-b91e-65274e5f4643 · outbound
Visual Lexicon: Rich Image Features in Language Space SciCap: Generating Captions for Scientific Figures
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77acda40-70fd-45e1-8a62-9ffb034197ec · outbound
Visual Lexicon: Rich Image Features in Language Space LoRA: Low-Rank Adaptation of Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 820a3d38-192c-4903-8b28-2b68487e3ded · outbound
Visual Lexicon: Rich Image Features in Language Space Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 214ad2ed-739b-401b-a78e-4a7e7ceb5abb · outbound
Visual Lexicon: Rich Image Features in Language Space Soda: Bottleneck diffusion models for representation learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d048583f-2d3a-4248-995e-33dbe42e710a · outbound
Visual Lexicon: Rich Image Features in Language Space Open- clip, 2021
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cbb5b517-af0f-4458-aca6-fa74ef836242 · outbound
Visual Lexicon: Rich Image Features in Language Space Referitgame: Referring to objects in pho- tographs of natural scenes
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 48d02701-c186-4f6d-a14f-6b27347adc43 · outbound
Visual Lexicon: Rich Image Features in Language Space Auto-Encoding Variational Bayes
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5722e5a2-58b1-469f-8f16-4016ff76d8ee · outbound
Visual Lexicon: Rich Image Features in Language Space Your diffusion model is secretly a zero-shot classifier
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60760f2a-1876-4691-81ba-2bfdefc76522 · outbound
Visual Lexicon: Rich Image Features in Language Space LLaVA-OneVision: Easy Visual Task Transfer
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f506868-c561-40dd-ba55-39aa22e4f875 · outbound
Visual Lexicon: Rich Image Features in Language Space Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2881563e-06ee-46ce-a4d5-84293a2247eb · outbound
Visual Lexicon: Rich Image Features in Language Space ImageFolder: Autoregressive Image Generation with Folded Tokens
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4234095f-c9a6-412e-920f-8b00f32cbcad · outbound
Visual Lexicon: Rich Image Features in Language Space Vila: On pre-training for visual language models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 51f159c9-6975-4607-9083-ce26e41d3649 · outbound
Visual Lexicon: Rich Image Features in Language Space Microsoft coco: Common objects in context
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 31ef5b21-af4d-4adb-af7c-7cf765f11b39 · outbound
Visual Lexicon: Rich Image Features in Language Space Improved baselines with visual instruction tuning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73a78549-8fda-415f-b5fd-31b0fe850442 · outbound
Visual Lexicon: Rich Image Features in Language Space Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 16f3383e-3957-4fb6-87bd-f7260ef7f4b9 · outbound
Visual Lexicon: Rich Image Features in Language Space Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9a4194e-ccde-4b59-b7a6-20b105456862 · outbound
Visual Lexicon: Rich Image Features in Language Space Prompting hard or hardly prompting: Prompt inversion for text-to-image diffusion models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2c298753-c6e2-42fc-aa86-fc2884543f2c · outbound
Visual Lexicon: Rich Image Features in Language Space Generation and comprehension of unambiguous object descriptions
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 11192204-bfe0-41cd-99de-3e2e4013ef88 · outbound
Visual Lexicon: Rich Image Features in Language Space Ok-vqa: A visual question answering 10 benchmark requiring external knowledge
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 343df608-9cd4-4d54-8b63-dabefd91fb01 · outbound
Visual Lexicon: Rich Image Features in Language Space Finite scalar quantization: Vq-vae made simple
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7e9a96d0-4db1-4a8f-85d1-d0e9c6846b71 · outbound
Visual Lexicon: Rich Image Features in Language Space Null-text inversion for editing real im- ages using guided diffusion models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25210b1e-3cda-4333-a301-8ae368bff820 · outbound
Visual Lexicon: Rich Image Features in Language Space Improved denoising diffusion probabilistic models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c3683f6d-6730-4d9f-b0d9-bcad9a0e70ad · outbound
Visual Lexicon: Rich Image Features in Language Space DINOv2: Learning Robust Visual Features without Supervision
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 780337ba-4573-47fb-973f-c07c08df00a6 · outbound
Visual Lexicon: Rich Image Features in Language Space Styleclip: Text-driven manipulation of stylegan imagery
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6972a36d-6ad5-4e07-979e-bd5d9eae5668 · outbound
Visual Lexicon: Rich Image Features in Language Space Sdxl: Improving latent diffusion models for high-resolution image synthesis
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3526c95c-494c-45a0-9b7d-392445279843 · outbound
Visual Lexicon: Rich Image Features in Language Space Learn- ing transferable visual models from natural language super- vision
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 07ccb4c7-7fe7-4226-9664-470eb4c50e37 · outbound
Visual Lexicon: Rich Image Features in Language Space Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 97de4469-4aca-463b-ac3c-1f848dbc3502 · outbound
Visual Lexicon: Rich Image Features in Language Space Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08351e34-873d-4549-b142-3807ea059427 · outbound
Visual Lexicon: Rich Image Features in Language Space High-resolution image syn- thesis with latent diffusion models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1741204b-4a84-47fc-8dc5-cca29e09855f · outbound
Visual Lexicon: Rich Image Features in Language Space U-net: Convolutional networks for biomedical image segmentation
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e9180746-7448-4671-93cd-2e7235e8e822 · outbound
Visual Lexicon: Rich Image Features in Language Space Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 726ffdd5-ad93-4285-875c-1b79fdb9ee05 · outbound
Visual Lexicon: Rich Image Features in Language Space Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b21586a3-038e-4654-ad79-7db984e8432f · outbound
Visual Lexicon: Rich Image Features in Language Space Photorealistic text-to-image diffusion models with deep language understanding
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9cb61f0b-e9c7-4677-80ee-4c886789ddef · outbound
Visual Lexicon: Rich Image Features in Language Space Progressive distillation for fast sampling of diffusion models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e1062836-6fdf-430a-b5fe-b1e2ceafd75c · outbound
Visual Lexicon: Rich Image Features in Language Space Neural Machine Translation of Rare Words with Subword Units
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbb4e182-7d1f-47d6-aba4-56cd33b86dd6 · outbound
Visual Lexicon: Rich Image Features in Language Space Adafactor: Adaptive learning rates with sublinear memory cost
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 98c570c5-8c58-4703-9090-7ceb5fd670a4 · outbound
Visual Lexicon: Rich Image Features in Language Space Textcaps: a dataset for image caption- ing with reading comprehension
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a0cd468-0675-418e-adb7-e63546c0b297 · outbound
Visual Lexicon: Rich Image Features in Language Space Towards vqa models that can read
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40168739-51f3-494d-82b2-52916c5d5d37 · outbound
Visual Lexicon: Rich Image Features in Language Space Deep unsupervised learning using nonequilibrium thermodynamics
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9d5d67f3-4ce2-4d40-8d65-5b72cb5ddac6 · outbound
Visual Lexicon: Rich Image Features in Language Space Generative modeling by esti- mating gradients of the data distribution
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c0291128-0b6e-4d13-b777-f912c63b328c · outbound
Visual Lexicon: Rich Image Features in Language Space Rethinking the inception archi- tecture for computer vision
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92366efc-2d27-4d38-a7f1-4ad0d2a9d062 · outbound
Visual Lexicon: Rich Image Features in Language Space Gemma: Open Models Based on Gemini Research and Technology
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72cfda0f-2a18-454e-90c0-ee95e2dbf7d5 · outbound
Visual Lexicon: Rich Image Features in Language Space Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6ef4604-a898-4f20-9585-20a277c9e7ea · outbound
Visual Lexicon: Rich Image Features in Language Space Visual autoregressive modeling: Scalable image generation via next-scale prediction
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3617e7a4-14ec-4c25-9ab0-df3055614c94 · outbound
Visual Lexicon: Rich Image Features in Language Space Con- trastive multiview coding
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 787f805e-c67a-4a2f-b217-0d5cea682b99 · outbound
Visual Lexicon: Rich Image Features in Language Space Extracting and composing robust features with denoising autoencoders
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bed01bb0-dc17-4af1-a093-924f36bfa87c · outbound
Visual Lexicon: Rich Image Features in Language Space Diffusion Feedback Helps CLIP See Better
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ca3aeca-eb7a-4e63-b017-670a2fa7dc70 · outbound
Visual Lexicon: Rich Image Features in Language Space Unsupervised feature learning by cross-level instance-group discrimina- tion
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cefa8917-1ddc-48fc-8ebb-d62875f10f97 · outbound
Visual Lexicon: Rich Image Features in Language Space De-diffusion makes text a strong cross- modal interface
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d6b9ab6d-99e3-45f6-abac-85d7c6cd926e · outbound
Visual Lexicon: Rich Image Features in Language Space Unsupervised feature learning via non-parametric instance discrimination
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ea83b458-1ff4-4fb3-aa0f-d4a9ebc34188 · outbound
Visual Lexicon: Rich Image Features in Language Space Msr-vtt: A large video description dataset for bridging video and language
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c8af8a4d-ca0b-4740-8d1b-0dc86c4cc6d6 · outbound
Visual Lexicon: Rich Image Features in Language Space Open-vocabulary panop- tic segmentation with text-to-image diffusion models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 13b41e26-cd54-491f-a928-9d7d24d32d7a · outbound
Visual Lexicon: Rich Image Features in Language Space Diffusion model as repre- sentation learner
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8ef7244e-c05c-46dc-8b6c-94b0e5886f8e · outbound
Visual Lexicon: Rich Image Features in Language Space Vector-quantized image modeling with improved vqgan
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cfc5324d-733c-4bbe-9044-d46b3037695a · outbound
Visual Lexicon: Rich Image Features in Language Space Coca: Contrastive captioners are image-text foundation models
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 17770810-7065-4cf0-9f7f-97014c3b3370 · outbound
Visual Lexicon: Rich Image Features in Language Space Modeling context in referring expres- sions
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cf55616b-e5eb-4651-afb4-d0794ba586d7 · outbound
Visual Lexicon: Rich Image Features in Language Space Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e45783fa-396d-4806-9a3c-bc9e234330df · outbound
Visual Lexicon: Rich Image Features in Language Space Language model beats diffusion–tokenizer is key to visual generation
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbda5d99-2584-4952-9df4-2b260fc56b0d · outbound
Visual Lexicon: Rich Image Features in Language Space An image is worth 32 tokens for reconstruction and generation
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9cfe70ec-22c6-428e-8c94-9b9fd7270e0a · outbound
Visual Lexicon: Rich Image Features in Language Space Soundstream: An end- to-end neural audio codec
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 69467f88-3b85-4b8c-ba8d-eb8f8942aaf0 · outbound
Visual Lexicon: Rich Image Features in Language Space Sigmoid loss for language image pre-training
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2e2c4c98-8c3b-434a-857d-ddd25aef3edd · outbound
Visual Lexicon: Rich Image Features in Language Space Image and Video Tokenization with Binary Spherical Quantization
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e599697c-7bea-411d-b4f2-60dcfa7a302b · inbound
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models Visual Lexicon: Rich Image Features in Language Space
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.