Pith. sign in

Paper Citation Record · LEDGER

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens

As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2506.03096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03096 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:14:28.606467Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:55:28.067052Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T11:55:29.193720Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved15
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ed39858-fade-4067-98e7-e06caba8e887 · outbound

This paper cites 4M-21: An any-to-any vision model for tens of tasks and modalities.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens 4M-21: An any-to-any vision model for tens of tasks and modalities

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:34.557399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:23.087022Z digest=sha256:9d8c9386e83f57519f1edd1195a63cea3880cc8888ca35d572936c05cb260504

Observation 9cb93a64-e857-4b93-9163-3b379ce22227 · outbound

This paper cites Zero-shot composed image retrieval with textual inversion.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Zero-shot composed image retrieval with textual inversion

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:34.296341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:23.178855Z digest=sha256:4a8c9dbd24b5cc7b55b46bbef96263585dbdecabb8ed76ac11282eb1fe55dfe7

Observation 843f6fe7-90e7-4cd0-9951-977d031bbeda · outbound

This paper cites BEiT: BERT pre-training of image transformers.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens BEiT: BERT pre-training of image transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:34.129683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:23.280121Z digest=sha256:4b8d39e3d7773a96a0c3059af7d261f44310de8568771de544c8772f97eac5b2

Observation 1b869257-70c1-4cce-97e8-a88289041117 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:23.390643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:23.390643Z digest=sha256:b80db2721c0cfb20d103ce455919b2fa7279dea75b0c2f53ea87b5867ec5fbaf

Observation d894bb48-436a-48e4-ba01-d7f94534e70f · outbound

This paper cites All you may need for vqa are image captions.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens All you may need for vqa are image captions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.990034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:23.517383Z digest=sha256:a83d03b6d50ae9511a455ca07e596cb07609ee3384cafc9e59895b63424dcde2

Observation 494f61e7-b2b6-47b9-9471-a78c51515c42 · outbound

This paper cites Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.886746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:23.580996Z digest=sha256:0a76dea8119d9727666a63d87c084876dfc68c15df3834b32329a532e68946f4

Observation 08e897f4-988d-438c-bc76-ef9b4c3b8f12 · outbound

This paper cites Understanding transferable representation learning and zero-shot transfer in CLIP.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Understanding transferable representation learning and zero-shot transfer in CLIP

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.775659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:23.639244Z digest=sha256:42d9a52cbef5f487d30ae9f6e03453bd6f38ddfaeaf0a15ce1ca44e059819d13

Observation 1cf73c11-e53e-4c18-9520-d7be9186633c · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Reproducible scaling laws for contrastive language-image learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.682817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:23.759334Z digest=sha256:491485a784a154f29bedb37c066e47183470f41fdea908429a3a22dbfc1e3330

Observation aa8602d3-d47b-443a-b5d2-dcfca48a8565 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Imagenet: A large-scale hierarchical image database

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.557702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:23.837170Z digest=sha256:092900206c73c8731af6dcca27a29d5919c5b326c679dd4819d9aad9530d7556

Observation cb770c8c-2a79-4e1d-adb2-3049e3c59a6a · outbound

This paper cites Bert: Pre-training of deep bidirec- tional transformers for language understanding.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Bert: Pre-training of deep bidirec- tional transformers for language understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.464984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:23.956817Z digest=sha256:9e8ac0bef49677504ee2eb1e701b4057c7f28f660167e981cfd6356e7c09a6fa

Observation c432fe62-658c-4d53-ad9a-a9022fc97142 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens An image is worth 16x16 words: Transformers for image recognition at scale

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.354101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:24.060445Z digest=sha256:7e0a2f1a9994955da651b29ff2842c95c942e1aeab380137e6770833874af022

Observation 0fd1af9d-39f7-4783-9788-f1d0f4d9fbdb · outbound

This paper cites The Llama 3 Herd of Models.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:24.193921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:24.193921Z digest=sha256:7dbd0d82a6b5745688464cc948e71d01cf0c112e139dac44b9caa7fb1ae2d10b

Observation b11d3bb8-0ab1-4fa2-8f9b-27fc5fc7f308 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Taming transformers for high-resolution image synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.245457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:24.311637Z digest=sha256:2f39e8ac23af12ed88699cc72b4e53f39d13568dbee877db2d9dd9104d8555ca

Observation bb475763-a59d-427a-b4c8-7ae4b4a5f9c4 · outbound

This paper cites Dreamsim: Learning new dimensions of human visual similarity using synthetic data.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Dreamsim: Learning new dimensions of human visual similarity using synthetic data

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.129797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:24.448475Z digest=sha256:9ecced73c87ac76aeb176ababb0eb9f11faa7392cb70bd058e33314bfe9829de

Observation 7e2d01de-0b28-4cb6-8211-104430daa9f5 · outbound

This paper cites Language-only efficient training of zero-shot composed image retrieval.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Language-only efficient training of zero-shot composed image retrieval

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.021382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:24.584746Z digest=sha256:c2bd843c6767da8a3cac8978157a4398475035764ae94f628f1cb64022edc1df

Observation 837e14ff-b602-40d5-b2fc-b2d9550d026a · outbound

This paper cites Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.919515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:24.674048Z digest=sha256:e3cd4c46b5efbc6a0544cd9ff8aba732a96ca26f4dd25c917261b7c15191f98a

Observation f5f16587-952d-4cf0-b754-d2c365832953 · outbound

This paper cites Hq-edit: A high-quality dataset for instruction-based image editing.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Hq-edit: A high-quality dataset for instruction-based image editing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.837374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:24.789096Z digest=sha256:d648ede10d6a630ded5f8bba704ad3e47b19600c014eb23dd3af94ed79ea061b

Observation df0adcd8-bfce-482c-a84c-7dae883e1ac1 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Scaling up visual and vision-language representation learning with noisy text supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:24.894088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:24.894088Z digest=sha256:b12f485a48716e376c8dd206a96996fa3c219dd02d1354f5b269b1282df5a8bd

Observation e334eac3-8f55-4ee9-8da4-99fb353502f9 · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:24.985554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:24.985554Z digest=sha256:c46cb2b876d72b75085cc0b9a56770f902872709fd0ef98b3e24741eb1679a7b

Observation b801fa92-b8c0-4041-8bb0-87e2ce385ef8 · outbound

This paper cites Vlm2vec: Training vision-language models for massive multimodal embedding tasks.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Vlm2vec: Training vision-language models for massive multimodal embedding tasks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.716035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:25.098746Z digest=sha256:19276e2bfe4562a3434fdbdb8577e38957569f52a00fcb8232aeefa428c5c7f7

Observation 5c55ce28-7e71-442b-9775-b612e8373450 · outbound

This paper cites Hard negative mixing for contrastive learning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Hard negative mixing for contrastive learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.613227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:25.212454Z digest=sha256:06733881b046c38642bc5b4a32f1a4d7fd51e1b09130e932c773aeaade98a69a

Observation b7f3d2fa-61a0-499e-b276-ddfe51cae974 · outbound

This paper cites Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:25.297748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:25.297748Z digest=sha256:413f6794753d5feafa3c7f2c0c045898b7c4df0c7ed0dc5f040ad2ee79642abd

Observation 163d1407-a624-4fcd-ba8f-b3bfae928cfc · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.IJCV, 2017.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Visual genome: Connecting language and vision using crowdsourced dense image annotations.IJCV, 2017

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.525438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:25.415308Z digest=sha256:00067e028e5e96cb2dfb8f52e5ab575f3138f670ebedb0cb0e146d040e203a1e

Observation 3d965c66-e85b-4216-a5ae-0e563dfeb644 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.417980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:25.518996Z digest=sha256:71c9e43effbfc8b5171b65cab8a95e4fc39ebcf96bab1cf76fe62700d8c633d0

Observation 7d89c8a0-66b4-4823-8681-f9057a169c61 · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.304284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:25.630094Z digest=sha256:5288ba87c303ff5288c2d4ec3411e9628383e25bbac170247f83d190cf70a597

Observation e07f0e3f-31d0-427a-8f86-cc262e0424ec · outbound

This paper cites Visual instruction tuning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Visual instruction tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:25.764193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:25.764193Z digest=sha256:2dfc65d2cf50bbeb9a2a1f7fe57f2d7bc1db1a0fd173177810dd7c3cc0bf472f

Observation 9267b3c3-057a-4a59-b17b-dc84985ae7ee · outbound

This paper cites Universal vision-language dense retrieval: Learning a unified representation space for multi-modal retrieval.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Universal vision-language dense retrieval: Learning a unified representation space for multi-modal retrieval

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.155301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:25.869859Z digest=sha256:5e230f515bd36f302bf03e5a91af854947d0a324043d7ced9013df4a0720cb07

Observation f831d799-616a-4c29-a57a-844c122c90c8 · outbound

This paper cites Decoupled weight decay regularization.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Decoupled weight decay regularization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.992247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:25.953221Z digest=sha256:ef88261508e8bb39d99d6ee7179b992674f8d272b831925f33bbeb9538416f50

Observation 0330e435-b54c-45b0-8640-3fb3636ff3be · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.091512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.091512Z digest=sha256:d01e6d3ee1e0dd27230a14ccbece06ad9467b13830ef4683191cf6684e412faa

Observation ebb90a6d-3e34-4faf-94fa-06e656d78717 · outbound

This paper cites 4M: Massively multimodal masked modeling.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens 4M: Massively multimodal masked modeling

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.867994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:26.267694Z digest=sha256:66ce13f51d8fb0df708951f257693f7c5eb2e6025defa9edc9f30e7dbb064f66

Observation 5ca6bf1b-f61f-42b0-9b94-3fb20cabc4f3 · outbound

This paper cites Learning transferable visual models from natural language supervision.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Learning transferable visual models from natural language supervision

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.778503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:26.367835Z digest=sha256:238ecd9d379a5806339f5bb9c75e1bc8adf0225783cffc822b0dbbb9e7fb7e16

Observation 685d25df-f3bd-4182-9146-cbfbb3299bdf · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.649413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:26.479501Z digest=sha256:d7a847d6893c7fdf4903614e4b5d87dd6d42dfa64226f59dda485999fa4f9063

Observation 726aa4b1-bdb6-4c82-bcfc-c29f883b32f8 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.584738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.584738Z digest=sha256:75a261fbec4af259d25141aa16bbe4e61540112e8f193abcdfffd840aa8f7a08

Observation fe737f5e-fc12-40b2-bf4a-a691b7b3fc69 · outbound

This paper cites Contrastive learning with hard negative samples.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Contrastive learning with hard negative samples

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.556519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:26.665589Z digest=sha256:bd21435b21ca284a7bc423d0ffa27b9a61f08cb43d3094ca76cd82ee9caf8ddb

Observation a161c2c6-7052-44f3-b35f-02ce37ba2616 · outbound

This paper cites Pic2word: Mapping pictures to words for zero-shot composed image retrieval.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Pic2word: Mapping pictures to words for zero-shot composed image retrieval

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.418550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:26.746652Z digest=sha256:e548084ff433845073d8c3ffe4e7a8ea0634cf1df9275cd79c593d6e08112cd9

Observation 0b3b8330-ed67-48e4-be40-aaa806d0b0c1 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.299067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:26.802857Z digest=sha256:0db7451b1e7f41fb942bb3c4381a87b719346d6faff3de52e49f2a1924479765

Observation e467b26d-a173-4250-b527-28093c29a0e8 · outbound

This paper cites Towards understanding the modality gap in clip.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Towards understanding the modality gap in clip

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.178468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:26.901538Z digest=sha256:db7528d1edc7039567ada9144f212624d0e9348e7fa5d887e62e90a2fd46d0fb

Observation 90f3f5d9-5e2a-4835-8190-6877d5b1b812 · outbound

This paper cites FLA V A: A foundational language and vision alignment model.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens FLA V A: A foundational language and vision alignment model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.046467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:26.965593Z digest=sha256:230fa460768dd9d8e1bba5734a2c7ce1676f38f753aa61647120ef15c6268d23

Observation a10a96e6-01a0-4da0-9c8d-1df34c23561d · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.042563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.042563Z digest=sha256:6decac2e86fa66612b9034a02abb78a7dda54b1fb78d70c207f9b160b380437f

Observation 613cdf8f-2a49-4ec4-8fa3-f44c17f4e834 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.136459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.136459Z digest=sha256:8d7e884ae01e0e71dbaa94135e4f23ad9f03df6e00df37c432ebab8fd3b8062c

Observation 497cf7b5-a8b4-43e8-8fea-914f796a061b · outbound

This paper cites Neural discrete representation learning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Neural discrete representation learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.251290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.251290Z digest=sha256:867b6bfeb90f439f39807da3eea8b26b4dba553970c6904d639545643934ac9b

Observation 05ee0f4a-8ca8-4684-bd1b-1a62196deba7 · outbound

This paper cites Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to- sequence learning framework.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to- sequence learning framework

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.869888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:27.350816Z digest=sha256:b33ed0b86a9dd97287603317ca865ef6e2e7ee871758148c86c1271d4b4a551d

Observation 505e623d-3c28-4808-bf86-62e062430458 · outbound

This paper cites Uniir: Training and benchmarking universal multimodal information retrievers.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Uniir: Training and benchmarking universal multimodal information retrievers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.710665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:27.482449Z digest=sha256:b221dac099fa6db27751ce74a8288088b4b7a40250fa38a5c28e8238f243bb0f

Observation 112ca790-d670-4692-8f41-5d4b31f18c51 · outbound

This paper cites Robust fine-tuning of zero-shot models.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Robust fine-tuning of zero-shot models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.570571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:27.604662Z digest=sha256:19aa06f1025db6b5281d5532db903aae420ecc3ea86c795be152dc9f9eda0d68

Observation 9d2c0216-9f12-4504-8b26-37d8d0e3c5f9 · outbound

This paper cites Bridgetower: Building bridges between encoders in vision-language representation learning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Bridgetower: Building bridges between encoders in vision-language representation learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.449384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:27.701150Z digest=sha256:fb5b3488201c7e7e805864617d4e95574d13c4a4dc421852926d6985a611573e

Observation f278eb64-6d5f-4f50-ab75-56c4c614c8fd · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens An image is worth 32 tokens for reconstruction and generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.310541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:27.794105Z digest=sha256:32876362b0a8e5ba03fb5b3d7197c810f3262d0685c518e5214b2362129eeb4c

Observation 01dc3f6d-aaf4-4c64-b5ba-ac51523ae568 · outbound

This paper cites Sigmoid loss for language image pre-training.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Sigmoid loss for language image pre-training

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.160191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:27.896708Z digest=sha256:1ef29b11ccbe9c4e56a8d5ca574b37de7d1b3137591faa845169f6c1038b06f3

Observation d637248f-3224-45c9-ab61-2d67cb6366af · outbound

This paper cites Magiclens: Self-supervised image retrieval with open-ended instructions.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Magiclens: Self-supervised image retrieval with open-ended instructions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.023795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:27.988902Z digest=sha256:94f0e98b49ac1771a7270d16dbb85174e1f45145018fd1516d1f3d7de97049f7

Observation d2b67434-2652-493e-b0bb-6da00b6f1b67 · outbound

This paper cites upper left, upper center,.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens upper left, upper center,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:29.864818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:28.101210Z digest=sha256:41bd141b0f7cc8d7d9a0a86500f8d25afecc98faa31483dc7ab2a9a02b93bee8

Observation 87177a54-1e14-4952-9ab1-f70f33b5a5aa · outbound

This paper cites an unresolved cited work.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:29.704991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:28.138032Z digest=sha256:e0274c4ba5701196af0da76899c520076768a5e9b3cf32015707af73cb85d487

Observation 484e56ad-241c-484d-8993-449bba093f24 · outbound

This paper cites an unresolved cited work.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:29.574570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:28.217502Z digest=sha256:cffbf5940cefc2690104c51bec96d197551c6faec97150b80a82a6b0ed37d463

Observation 7691d229-b6cf-4efc-8da9-85f6e5f3b099 · outbound

This paper cites an unresolved cited work.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:29.421030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:28.321253Z digest=sha256:894948900abad5b3e36dd5a1805c82729c62bc38b2ce4efc50d8e9fd6a2ae88d

Observation dc538a3b-e078-43bc-9b44-b08d22f7d502 · outbound

This paper cites an unresolved cited work.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:29.252149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:28.404311Z digest=sha256:7028e1029ecf199788e13dd9d0610ff4bccf6cb492f2a2b8fa72638b387d7b8c

Observation dd478e66-6fbf-4a1c-81b5-6c79b2bb817e · outbound

This paper cites The {object_name} on the left/right.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens The {object_name} on the left/right

Reference 54

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:14:29.054862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:28.511265Z digest=sha256:1ff5b76f4e792bf551b6a642f5077e011cc86221b29e22825a802fe5fa6ee2d8

Observation 53a2cc21-01d7-4499-973b-e6ec13121e93 · outbound

This paper cites Notably, this model is much larger in the amount of parameters (4.15B, i.e.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Notably, this model is much larger in the amount of parameters (4.15B, i.e

Reference 55

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:14:28.889029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:14:28.606467Z digest=sha256:428c417d009d4475680b9a2495c722715f039307b2b6dbd633d46cfb51f5494e

Pith citing papers

Observation cdce7db5-aaeb-44ed-b0dc-df797bdead84 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:55:29.196787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:55:28.067052Z digest=sha256:90091b0ff128ee8f50df4a3eeca5ecc463ca930d4f156fc002ed0564d72361ae