Pith. sign in

Paper Citation Record · LEDGER

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

As of 17 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 13 inbound Pith citation observations for arXiv:2506.23115.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23115 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:52:26.368534Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:30:14.527131Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:26:45.637665Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved43
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 12acaddd-f71c-4817-9d01-a3f52b1362f7 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:21.406024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:21.406024Z digest=sha256:15433ac0115b886222e6fe211d6838c144741a3fc16dd116630c25170507d90e

Observation 82d96836-307e-4b52-9786-ff31030520a3 · outbound

This paper cites Qwen2.5-VL Technical Report.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:21.501578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:21.501578Z digest=sha256:aa0225cae797e23638c44511f5244be78f86e78d75871c7c6e58d317592b84f8

Observation 909fba20-31a7-404e-aec4-813509e6093f · outbound

This paper cites Vlmo: Unified vision- language pre-training with mixture-of-modality-experts.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Vlmo: Unified vision- language pre-training with mixture-of-modality-experts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:21.581303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:21.581303Z digest=sha256:542219fa5d34ea0499308217aa9e2c67a9898e4464276afbe1a260208c8e61fb

Observation 435d780e-2ef9-4dae-8ff7-148ce382aeef · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:21.664520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:21.664520Z digest=sha256:009b78ebff27665e62ae2bd118c164aa8417fab7a0743c27939d41c7a2fb5083

Observation 8ef96816-7451-403d-b809-34d0de0c171d · outbound

This paper cites mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:21.757887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:21.757887Z digest=sha256:8f0c50a414cdc35d4e3437077355f33c1ea128734b82fc16a59d3699528d0aac

Observation 6bb383f2-3ef2-4592-b394-8b50ba4b9ab0 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:21.862210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:21.862210Z digest=sha256:ca92dacbc0290dee1620d613ea1a94e8e5898cc8866cf0dca11e27c43bbb0a87

Observation d6831270-7664-4226-a823-c87311c8f726 · outbound

This paper cites UNITER: universal image-text representation learning.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings UNITER: universal image-text representation learning

Reference 7

Resolution
malformed identifier
no resolver link, observed 2026-08-06T21:52:21.927001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:21.927001Z digest=sha256:d2b2bd670ea8f242d70298b722705502c0ddce4d40afa10dde94636014f8c9b8

Observation ba50800f-6021-4c9a-b239-f00b9671e97f · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Reproducible scaling laws for contrastive language-image learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:22.042338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:22.042338Z digest=sha256:a874dae17c33f1b4e6bb43124dc12f940e337409e3313a4b7b02408a0295aca6

Observation 74616016-8654-4317-99d1-9ff2b53422f4 · outbound

This paper cites BERT: pre-training of deep bidirectional transformers for language understanding.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings BERT: pre-training of deep bidirectional transformers for language understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.802175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:52:22.118312Z digest=sha256:a53d296d2149252ebb2405207003b15166219458894fac0a456069bc5be10456

Observation 8d2775c9-e589-4a2c-9b28-fad78c05067b · outbound

This paper cites Colpali: Efficient document retrieval with vision language models.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Colpali: Efficient document retrieval with vision language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.661796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:52:22.378707Z digest=sha256:481925edc66e1d728174e6fad0434ce9e2e962162dc7379a2c1fee24ca564936

Observation 2cb5873b-4130-4d15-a942-5302520d6146 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:22.484484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:22.484484Z digest=sha256:100be9204a42acc66ade6140a83ff3200991e4ac603e9de24d950fb364da0535

Observation b1c4f88b-a3a4-4dea-b468-f8ca7bdfed86 · outbound

This paper cites Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:22.602494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:22.602494Z digest=sha256:c9970e40d834b70b27a5ae2c0864731de3fae2b3cc480777dd6b81b056c14524

Observation de573936-c557-4175-a5a9-a506d29443bb · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:22.687226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:22.687226Z digest=sha256:54f355f98bb05636bf557590e1f14468cb4e01a600ed7779c70fb1ae5971b988

Observation e154328b-0d4a-4b0d-b4bc-2725ca9bb854 · outbound

This paper cites Girshick.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Girshick

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:22.796886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:22.796886Z digest=sha256:a22466a5569dc7ea226b508a741ca428fbb33e28ba95c27f9d7b1f11fe723a1b

Observation a6dbd464-fc12-49f0-9ff6-258cd273f447 · outbound

This paper cites Le, Yun- Hsuan Sung, Zhen Li, and Tom Duerig.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Le, Yun- Hsuan Sung, Zhen Li, and Tom Duerig

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:22.908229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:22.908229Z digest=sha256:c97567d778a70f065aca6b6633c9d39dc334c651f464c4c6f3bd2bdaf506807e

Observation f8732002-6067-48ac-8ef8-632b373dee34 · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.138389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.138389Z digest=sha256:4f8d43d0a647006fa5177edc4eddb3613f638d5efa6238c52e4722233c01065a

Observation 67d2c4ad-3ef7-497a-9315-4477152a8d1f · outbound

This paper cites Continual pre-training of language models.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Continual pre-training of language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.479218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:52:23.199431Z digest=sha256:62d929cc2238bcc184916af5860787c1ad8c778ce8aa4e59421dcc4d679c1b18

Observation 54ae3e34-60d7-4074-a352-4bab87dc0a52 · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Vilt: Vision-and-language transformer without convolution or region supervision

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.295768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:52:23.291964Z digest=sha256:474fb5e33c21cd692f3926eeab3682c35eda87f3bd510071a6e67905d6ca720f

Observation 43747104-f690-45a4-bb11-b893ba5279eb · outbound

This paper cites Llave: Large language and vision embedding models with hardness-weighted contrastive learning.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Llave: Large language and vision embedding models with hardness-weighted contrastive learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.463875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.463875Z digest=sha256:3a02fa1b4a07c2fb04551578bf27f8c396418ef58bef299a83f843501eec1ba7

Observation 2021fc53-9986-495d-b081-c67fdc160c3b · outbound

This paper cites Building and better understanding vision-language models: insights and future directions.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Building and better understanding vision-language models: insights and future directions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.673554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.673554Z digest=sha256:eb4d34a2e1a7922646fc6008235629ff00ba3140277c0f59c89e25aabbd731fe

Observation 2cb4484a-1647-4fc6-96ed-38e9d8b51b7c · outbound

This paper cites an unresolved cited work.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:30.115856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:52:23.382898Z digest=sha256:cbfbbf18377ed45317a94fc98922c2e5e7df45142f41d77445bc61a246972884

Observation e88daa49-3aaa-4164-9451-0c643901b11c · outbound

This paper cites an unresolved cited work.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:29.815446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:52:23.942267Z digest=sha256:cafa4d28edf72b60faa083baab9ae7450c6fafe841c3ebe47eb26467b394f173

Observation 2fc0f898-7539-423c-9f07-a3f2c902346f · outbound

This paper cites an unresolved cited work.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:29.616930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:52:24.064803Z digest=sha256:bb838e29dca5313063dbc124e9dc45206f7362e4c99f266c439181791f5c7ef1

Observation 27326480-7e37-4820-8117-a9e3fe2741b6 · outbound

This paper cites an unresolved cited work.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:29.440599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:52:24.181614Z digest=sha256:df112a986d843194d0d08174a5d88cfbf948593217e9ca26e6c8ba166dba0bb0

Observation a198e8bd-8db9-432a-8644-fe2d8b5e600b · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.269954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.269954Z digest=sha256:f25c74474796573d5251d36afc8decc09a5ec06f534aa1e2883d05b4d979beb6

Observation 467cab76-10a2-4460-b58f-003ed0a09128 · outbound

This paper cites Nv-embed: Improved techniques for training llms as generalist embedding models.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Nv-embed: Improved techniques for training llms as generalist embedding models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:29.950680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:52:23.851285Z digest=sha256:0828aef4a0304d8ea0e8d04004d0c2e4262dd2074c6f823d75cca67fde749819

Observation 04a6f182-78a7-4c78-bd37-b931664633a0 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Improved baselines with visual instruction tuning, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.483412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.483412Z digest=sha256:969b4f056ee399ec287221799385c32374efde3b9653ca27fc39b3d47a2aa3d0

Observation 40dacb9d-03dc-4323-9b56-d3b19b99167b · outbound

This paper cites Unify- ing multimodal retrieval via document screenshot embedding.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Unify- ing multimodal retrieval via document screenshot embedding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:29.257243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:52:24.556343Z digest=sha256:498719288de8f9140567a1b19f418c3393307b61f21ce495f8e6bf799624c669

Observation de5e1ea6-d8e2-4010-b025-6c62838ceed3 · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.627791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.627791Z digest=sha256:efc6e4b7d91b53ff515c966c3b0bf86071e2c8bc15b035de537b51d977888b72

Observation e4a46cba-2329-4e31-91f6-92d767521cf4 · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date, April 2024.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Introducing meta llama 3: The most capable openly available llm to date, April 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:29.007803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:52:24.672117Z digest=sha256:4308aff21dcc6ee0722cc5b7d1c143ca479d1a2b039be8d0b9eea717a79b97c3

Observation 27f972d3-bcaf-4402-ac69-6c91923ce885 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Learning transferable visual models from natural language supervi- sion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.721367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.721367Z digest=sha256:34210e5435cbb7db9cbc74eaec9cc013e7b5a9356018079a61c429a02379c0d3

Observation b54f5c2c-b25a-4369-8aeb-aafe3724a387 · outbound

This paper cites MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.419284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.419284Z digest=sha256:33db57e4099b17da130055ca724f8f4213ff90b997fc852c5fe3362a1e119890

Observation 33f331cd-9b92-4920-8f2d-4923efaa5809 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.853435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.853435Z digest=sha256:10769158652ee6a420f9b7790215b117432ab5a07327373c5c09cffefb5a44e7

Observation 4aaaa547-1faf-4365-b5ba-e27ee147e874 · outbound

This paper cites From Pixels to Prose: A Large Dataset of Dense Image Captions.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.893267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.893267Z digest=sha256:bf3cc37f1de981f74cd6f74b947ed9091097ded8fae5b1d326f19bcf095a438e

Observation 9a1d01eb-cbd1-4ecf-bc7a-68583e134752 · outbound

This paper cites Repetition improves language model embeddings.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Repetition improves language model embeddings

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:28.720479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:52:24.962392Z digest=sha256:1d2f0e5c7a49bc02cfbed4c351404ebcc157bd7e2a466b49c08a2d7771c86b2b

Observation c75817ac-a642-436b-97f9-5059288eb293 · outbound

This paper cites LXMERT: learning cross-modality encoder representations from transformers.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings LXMERT: learning cross-modality encoder representations from transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:28.497621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:52:25.018644Z digest=sha256:3f55ceda7168f81737a3dc7c2cc8f2d3395bd53e10122d55cf7d4b965e6282f1

Observation 86f98179-3dd5-419a-85b2-f97c92455e12 · outbound

This paper cites BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.118584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.118584Z digest=sha256:586c11ad64fd22ee1afa2631fa755213ad6b8b1bf774c8e92b7ad58550697363

Observation 5bafc57e-5de7-4b00-b17c-64d39f540fec · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.800410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.800410Z digest=sha256:db96547e42bc18a356c1a8e53ea2ae8b5d92d2cf3c926b527188726dcb4f59e3

Observation 0f2bbf62-cfd6-4313-b2fd-2552ff1808fc · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.252459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.252459Z digest=sha256:dc8239f9e9c9d05668c375d55c7adeee7b9245a11bea80369fbd6d10d5516ee9

Observation 17a275a2-6222-4a70-afc4-b1203795b1cb · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.314513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.314513Z digest=sha256:9448d6c136b41930934e74addf33212463ca2a50e5847f640dc05b443461497d

Observation 751f62b2-c372-41a7-aa00-ceaba5a439fe · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Emu3: Next-Token Prediction is All You Need

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.380540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.380540Z digest=sha256:db53e836412bc29bc8a46e637f2a192c221a2ea1493977d5720324e4f6418add

Observation 1f4645f2-7322-4014-8f1c-438cecd41c9a · outbound

This paper cites Uniir: Training and benchmarking universal multimodal information retrievers.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Uniir: Training and benchmarking universal multimodal information retrievers

Reference 42

Resolution
malformed identifier
no resolver link, observed 2026-08-06T21:52:25.446061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.446061Z digest=sha256:210b2f9852fa319b5cbb898ae804314e37e33f78fa7e33ed2d98f2a6af414d7b

Observation c1c72e75-c1c5-4c1c-a89c-8b1a88ab05f9 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.514246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.514246Z digest=sha256:2206d9dabc45d8fb737abf018be2bb3ef5b97c9437894f7457d2b554144977f9

Observation 4a860906-83a2-4e4b-b708-f105f1a163e1 · outbound

This paper cites C-pack: Packed resources for general chinese embeddings.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings C-pack: Packed resources for general chinese embeddings

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.582577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.582577Z digest=sha256:ce9c63cd6988add2e5e96d521feb74457a615c8722e2212d33240655b6620a69

Observation cb2ad355-3c09-430c-a8e1-a6584a9659b6 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.203463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.203463Z digest=sha256:33752a76ab50a6ffc8f28fd2e3584119fdfb8bee57172725c9d7d5d911af5d00

Observation a21ad657-6990-46c3-84f9-2d33bc8c9f1b · outbound

This paper cites Visrag: Vision-based retrieval-augmented generation on multi-modality documents.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Visrag: Vision-based retrieval-augmented generation on multi-modality documents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.702492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.702492Z digest=sha256:2bf23beeca1838557215f562b79c6c02850840367da4f24523f89108f8ac1616

Observation 5bcfa3ea-7f46-49e2-a258-dc5fa0a2f259 · outbound

This paper cites Sigmoid loss for language image pre-training.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Sigmoid loss for language image pre-training

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.706693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.706693Z digest=sha256:edb72295532d6402521eb49146bb2ecce993ca67022f33c5328fe78bd46e2829

Observation e487b7ce-63d3-43cf-9f82-aaed71d2ba84 · outbound

This paper cites Magiclens: Self-supervised image retrieval with open-ended instructions.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Magiclens: Self-supervised image retrieval with open-ended instructions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:27.972882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:52:25.778184Z digest=sha256:83c4fe0ab15a3d9b2953699929c0c1d3c62cbd7766ed60a39566393630dc8523

Observation 38281db6-a1d0-4a2c-9c01-50fb9d8760fe · outbound

This paper cites GME: Improving Universal Multimodal Retrieval by Multimodal LLMs.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.999631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.999631Z digest=sha256:fdc2705f2a8df9515f5e28fee5a4123baef3330c8664a17f1106a0a8656632e7

Observation e389c8ad-e413-4311-96c6-dedb6e571120 · outbound

This paper cites QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:26.204540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:26.204540Z digest=sha256:22639ad3a02b1f5d2ddf1ad34323d2656b7868c136058b573d458f9f10304971

Observation ea48ff13-bf4e-40f4-a2a4-63577caebc06 · outbound

This paper cites MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

Reference 51

Resolution
malformed identifier
no resolver link, observed 2026-08-06T21:52:26.368534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:26.368534Z digest=sha256:56fdbe06c93d79f6278f389004491ef32ec457978231af273991e3390a076621

Observation 2f62bcb9-a8b1-4ee3-ba58-d06d0b7c1ae4 · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings An image is worth 32 tokens for reconstruction and generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:28.277738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:52:25.635632Z digest=sha256:7d9878185f2254764bf901eaffac83129493ce63a5d1b6cbc3a8348ec58cb488

Observation 567955fb-40b2-4acc-930a-47ea0fb70693 · outbound

This paper cites URL https://doi.org/10.18653/v1/D19-1514.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings URL https://doi.org/10.18653/v1/D19-1514

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.064163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.064163Z digest=sha256:28cd7505efc759175cdd62875ff367cbaa5999a975c08c1ceed9ce72b6163687

Observation ecbc4db1-4b25-4932-81d6-0100114c236c · outbound

This paper cites an unresolved cited work.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Unresolved cited work

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.026072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.026072Z digest=sha256:2015c9dfcdf1ee29892dc95b693bfd3360c90153bbf62f8c5ad02e5b6ce4efd6

Observation 79d78f86-b5d3-4f1a-9513-935d111541e7 · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.355520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.355520Z digest=sha256:bca96fc88622f204adf3c145fb951fdef7abab09319607ef3ded2ee2e57905e0

Observation de154433-0b74-42e1-91f1-83e5f4c6ba35 · outbound

This paper cites Building and better understanding vision-language models: insights and future directions.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Building and better understanding vision-language models: insights and future directions

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.743202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.743202Z digest=sha256:63bcb88150f5064eded4a3ae4c1af175e62b5dbe35438b6e76d3073682c6878f

Observation e3d2d85c-1d26-4dfb-95ab-6f1377d6f745 · outbound

This paper cites URL https://doi.org/10.48550/arXiv.2503.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings URL https://doi.org/10.48550/arXiv.2503

Reference 2025

Resolution
verified exact
doi, observed 2026-08-06T21:52:27.096202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:52:23.579127Z digest=sha256:c27b7185f253e31b07966f93a7ebbfadcaa184475326f159c0ba38e19c05ee84

Observation 338e1be0-732f-4800-ad45-472d56d5bf8a · outbound

This paper cites doi: 10.18653/V1/N19-1423.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings doi: 10.18653/V1/N19-1423

Reference 4186

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:22.242467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:22.242467Z digest=sha256:d3404ec20c477367dd61e0f7df9b468cd644094bed6e19da8587b06c0e7f56a9

Pith citing papers

Observation 4040c3b5-7f43-4da2-b596-f0cd0b882796 · inbound

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction cites this paper.

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:11:27.483851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-18T14:09:22.942238Z digest=sha256:70ab4037034ee14ef3be285275aa67851912b5920afaa63a3bc86ca254ffd685

Observation 12d175b6-6b6f-42b9-9584-86eab0bb3b58 · inbound

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories cites this paper.

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T01:02:32.688658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:02:32.688658Z digest=sha256:907b3b13620143509c6f5567aa3b178a27186111df89b45095dc64f0ca21d1aa

Observation 4604f3a0-ff6b-4c15-b94e-32d5feffcd74 · inbound

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning cites this paper.

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T19:41:31.172392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:41:31.172392Z digest=sha256:c563e5a0e0378936f528219c9c5c602bfb7fd6161c8a15c93da75438518e5e55

Observation 2346178e-b092-41c7-a44a-ce6802d66eb7 · inbound

PLUME: Latent Reasoning Based Universal Multimodal Embedding cites this paper.

PLUME: Latent Reasoning Based Universal Multimodal Embedding MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:19.962706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:cc0e7655910839e938205128dd3312c11cc3385a55ae0938e5557f7e9c1a87e5

Observation 155bb17e-45e4-4b13-bf9e-465976c22e96 · inbound

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings cites this paper.

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:01:19.810359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T12:46:30.827346Z digest=sha256:2a4842bf689363e35db6babcc2087c14120be8718ee3ba722849343212188e19

Observation 3f1d2f26-f696-4c90-9b78-03891cb2e354 · inbound

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings cites this paper.

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T18:31:37.144968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:31:37.144968Z digest=sha256:6de33f1cbdd9fd86ad3fa295760f52458a9025e2e73515a6044dcbb5b761dbdd

Observation 6e2d8627-e61f-471e-8fb1-880581564494 · inbound

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings cites this paper.

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T15:40:29.103187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:40:29.103187Z digest=sha256:82f258bf5d4667ef7f1af11e26720fc9fdfaf729f49e8fe03b3a3b28c7a33388

Observation 4d7edc2b-ef31-4c3e-b2c7-0e5a977e11f5 · inbound

MINER: Mining Multimodal Internal Representation for Efficient Retrieval cites this paper.

MINER: Mining Multimodal Internal Representation for Efficient Retrieval MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:06:10.555032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T12:40:33.437364Z digest=sha256:2dbcd51cc12d1ad5c721790e7fa005d6df82620ca682b8522da57b4aa912f4bd

Observation df16cd08-1747-4e39-93d3-9260da89c2f9 · inbound

FLUID: From Ephemeral IDs to Multimodal Semantic Codes for Industrial-Scale Livestreaming Recommendation cites this paper.

FLUID: From Ephemeral IDs to Multimodal Semantic Codes for Industrial-Scale Livestreaming Recommendation MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:21:16.786986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T08:21:06.959344Z digest=sha256:545ea284aaee5f447f71631b32a7fa91a176f3afb2b6e478fb674df6184d97c7

Observation 175ad642-69e0-43ee-9928-e0b288c9fc5a · inbound

FLUID: From Ephemeral IDs to Multimodal Semantic Codes for Industrial-Scale Livestreaming Recommendation cites this paper.

FLUID: From Ephemeral IDs to Multimodal Semantic Codes for Industrial-Scale Livestreaming Recommendation MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:54:59.172275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T16:45:43.932116Z digest=sha256:34525f53d09f38f06c7af792aa5fe62ff1a6052fbf97fbbdccdd38cecfb70aff

Observation 7ef45a18-3a16-4182-b019-468cb531e09f · inbound

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini cites this paper.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.317846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:0dbcf8928ac9fa4f34f62a18bf63d4aad88272dbbcbd7377bd222c6e76573c84

Observation f86d2da2-9aa9-42e3-b9be-d9577ab89c77 · inbound

MM-Matryoshka: Towards Budget-Elastic Visual Document Retrieval via a 2D Multimodal Matryoshka Training Framework cites this paper.

MM-Matryoshka: Towards Budget-Elastic Visual Document Retrieval via a 2D Multimodal Matryoshka Training Framework MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:26:45.639136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T06:59:06.801340Z digest=sha256:2ef1e34c5710731de8837ed06aa4d75ac342f44ef66a994e42aea7f6201156df

Observation 7f1dddab-eb44-4019-8c3f-db21b1a3070f · inbound

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval cites this paper.

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T18:30:14.527131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:30:14.527131Z digest=sha256:254e0a66296a1fb56bbb5825b99240503a282fd8e00fe7abdc794eff5349646a