Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:14:28.606467Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2506.03096.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:14:28.606467Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:55:28.067052Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T11:55:29.193720Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9ed39858-fade-4067-98e7-e06caba8e887 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens 4M-21: An any-to-any vision model for tens of tasks and modalities
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9cb93a64-e857-4b93-9163-3b379ce22227 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Zero-shot composed image retrieval with textual inversion
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 843f6fe7-90e7-4cd0-9951-977d031bbeda · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens BEiT: BERT pre-training of image transformers
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b869257-70c1-4cce-97e8-a88289041117 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens PaliGemma: A versatile 3B VLM for transfer
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d894bb48-436a-48e4-ba01-d7f94534e70f · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens All you may need for vqa are image captions
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 494f61e7-b2b6-47b9-9471-a78c51515c42 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08e897f4-988d-438c-bc76-ef9b4c3b8f12 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Understanding transferable representation learning and zero-shot transfer in CLIP
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1cf73c11-e53e-4c18-9520-d7be9186633c · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Reproducible scaling laws for contrastive language-image learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa8602d3-d47b-443a-b5d2-dcfca48a8565 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Imagenet: A large-scale hierarchical image database
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb770c8c-2a79-4e1d-adb2-3049e3c59a6a · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Bert: Pre-training of deep bidirec- tional transformers for language understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c432fe62-658c-4d53-ad9a-a9022fc97142 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens An image is worth 16x16 words: Transformers for image recognition at scale
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0fd1af9d-39f7-4783-9788-f1d0f4d9fbdb · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b11d3bb8-0ab1-4fa2-8f9b-27fc5fc7f308 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Taming transformers for high-resolution image synthesis
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bb475763-a59d-427a-b4c8-7ae4b4a5f9c4 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Dreamsim: Learning new dimensions of human visual similarity using synthetic data
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e2d01de-0b28-4cb6-8211-104430daa9f5 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Language-only efficient training of zero-shot composed image retrieval
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 837e14ff-b602-40d5-b2fc-b2d9550d026a · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f5f16587-952d-4cf0-b754-d2c365832953 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Hq-edit: A high-quality dataset for instruction-based image editing
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df0adcd8-bfce-482c-a84c-7dae883e1ac1 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Scaling up visual and vision-language representation learning with noisy text supervision
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e334eac3-8f55-4ee9-8da4-99fb353502f9 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens E5-V: Universal Embeddings with Multimodal Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b801fa92-b8c0-4041-8bb0-87e2ce385ef8 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Vlm2vec: Training vision-language models for massive multimodal embedding tasks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c55ce28-7e71-442b-9775-b612e8373450 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Hard negative mixing for contrastive learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7f3d2fa-61a0-499e-b276-ddfe51cae974 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 163d1407-a624-4fcd-ba8f-b3bfae928cfc · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Visual genome: Connecting language and vision using crowdsourced dense image annotations.IJCV, 2017
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d965c66-e85b-4216-a5ae-0e563dfeb644 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d89c8a0-66b4-4823-8681-f9057a169c61 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e07f0e3f-31d0-427a-8f86-cc262e0424ec · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Visual instruction tuning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9267b3c3-057a-4a59-b17b-dc84985ae7ee · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Universal vision-language dense retrieval: Learning a unified representation space for multi-modal retrieval
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f831d799-616a-4c29-a57a-844c122c90c8 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Decoupled weight decay regularization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0330e435-b54c-45b0-8640-3fb3636ff3be · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebb90a6d-3e34-4faf-94fa-06e656d78717 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens 4M: Massively multimodal masked modeling
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ca6bf1b-f61f-42b0-9b94-3fb20cabc4f3 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Learning transferable visual models from natural language supervision
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 685d25df-f3bd-4182-9146-cbfbb3299bdf · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 726aa4b1-bdb6-4c82-bcfc-c29f883b32f8 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe737f5e-fc12-40b2-bf4a-a691b7b3fc69 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Contrastive learning with hard negative samples
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a161c2c6-7052-44f3-b35f-02ce37ba2616 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Pic2word: Mapping pictures to words for zero-shot composed image retrieval
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b3b8330-ed67-48e4-be40-aaa806d0b0c1 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e467b26d-a173-4250-b527-28093c29a0e8 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Towards understanding the modality gap in clip
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90f3f5d9-5e2a-4835-8190-6877d5b1b812 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens FLA V A: A foundational language and vision alignment model
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a10a96e6-01a0-4da0-9c8d-1df34c23561d · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 613cdf8f-2a49-4ec4-8fa3-f44c17f4e834 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 497cf7b5-a8b4-43e8-8fea-914f796a061b · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Neural discrete representation learning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05ee0f4a-8ca8-4684-bd1b-1a62196deba7 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to- sequence learning framework
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 505e623d-3c28-4808-bf86-62e062430458 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Uniir: Training and benchmarking universal multimodal information retrievers
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 112ca790-d670-4692-8f41-5d4b31f18c51 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Robust fine-tuning of zero-shot models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d2c0216-9f12-4504-8b26-37d8d0e3c5f9 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Bridgetower: Building bridges between encoders in vision-language representation learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f278eb64-6d5f-4f50-ab75-56c4c614c8fd · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens An image is worth 32 tokens for reconstruction and generation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01dc3f6d-aaf4-4c64-b5ba-ac51523ae568 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Sigmoid loss for language image pre-training
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d637248f-3224-45c9-ab61-2d67cb6366af · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Magiclens: Self-supervised image retrieval with open-ended instructions
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d2b67434-2652-493e-b0bb-6da00b6f1b67 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens upper left, upper center,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87177a54-1e14-4952-9ab1-f70f33b5a5aa · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 484e56ad-241c-484d-8993-449bba093f24 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7691d229-b6cf-4efc-8da9-85f6e5f3b099 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc538a3b-e078-43bc-9b44-b08d22f7d502 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd478e66-6fbf-4a1c-81b5-6c79b2bb817e · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens The {object_name} on the left/right
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53a2cc21-01d7-4499-973b-e6ec13121e93 · outbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Notably, this model is much larger in the amount of parameters (4.15B, i.e
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cdce7db5-aaeb-44ed-b0dc-df797bdead84 · inbound
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.