Pith. sign in

Paper Citation Record · LEDGER

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval

As of 8 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2506.06144.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06144 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:05:31.866869Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:12:07.548447Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T23:20:53.347635Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19dd6684-388b-4cb7-be7d-caf1c71fb844 · outbound

This paper cites Qwen2.5-vl technical report,.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Qwen2.5-vl technical report,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.634670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.634670Z digest=sha256:7ab1005d49665551e2a11d43b0862c1158e963af6887da4671426686e357669e

Observation 9c47f90b-1e70-4dbe-9359-ec944a84271d · outbound

This paper cites Vlmo: Unified vision- language pre-training with mixture-of-modality-experts.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Vlmo: Unified vision- language pre-training with mixture-of-modality-experts

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.732933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.644806Z digest=sha256:dada9e2d019f53239621b7dafbe2e67796b215a68369d12c284b776a5122eaef

Observation 4d14f39b-27d1-4e85-af09-9a0b17dcb282 · outbound

This paper cites Valor: Vision-audio-language omni-perception pretraining model and dataset, 2023.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Valor: Vision-audio-language omni-perception pretraining model and dataset, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.713163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.652579Z digest=sha256:aa9e2457b23b98ff7108bc8a02a7f54ff6b6a34e3915049204663a18c8ef3162

Observation 204d3d2f-313a-435e-bdbf-f9d933524ab9 · outbound

This paper cites V AST: A vision-audio-subtitle-text omni-modality foundation model and dataset.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval V AST: A vision-audio-subtitle-text omni-modality foundation model and dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.702081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.656440Z digest=sha256:65d94715e69d04e5d6d13cedb2af5999ac0af9ca3554da5eebaf638b2a6de669

Observation 4e6e1dea-47dc-4f43-9209-8bdd96927bd8 · outbound

This paper cites PaLI: A jointly-scaled multilingual language-image model.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval PaLI: A jointly-scaled multilingual language-image model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.691985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.661364Z digest=sha256:f5441acc4b30893455164c3e6792b724bc0621b54cb9dbaa94e6c29f11305e4f

Observation 4041aeb2-69d3-403c-9bbd-41e128cb652d · outbound

This paper cites M3docrag: Multi- modal retrieval is what you need for multi-page multi-document understanding, 2024.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval M3docrag: Multi- modal retrieval is what you need for multi-page multi-document understanding, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.682596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.666238Z digest=sha256:a7887855e61db7c507ff2ef4fabdf307f743b16412d425428c76d095a8aa8da6

Observation c196a0ea-c869-4dbc-8144-f7bd9a7b07ff · outbound

This paper cites Jacolbertv2.5: Optimising multi-vector retrievers to create state-of-the-art japanese retrievers with constrained resources, 2024.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Jacolbertv2.5: Optimising multi-vector retrievers to create state-of-the-art japanese retrievers with constrained resources, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.672817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.675998Z digest=sha256:aad7fecb4dd89d40f9303714199f41dda4dfd38feb6970cace96850d4284c2a3

Observation 823aed3b-6806-41b2-87b7-e11f5308cda9 · outbound

This paper cites Cormack, Charles L A Clarke, and Stefan Buettcher.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Cormack, Charles L A Clarke, and Stefan Buettcher

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.680329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.680329Z digest=sha256:6338d1e11e87d561466424aad96dd4c436560c419627531a929f39c612861da7

Observation fdb88050-ddff-42c2-a84a-9587fc3bd8b6 · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval QLoRA: Efficient Finetuning of Quantized LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.684465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.684465Z digest=sha256:a58cc1cd500dc6d2b7d355d9698f834903aa9b726f9bf55afdb3b53378afdc95

Observation e94752d0-2164-4e0b-b8ee-8b410a9fdcc2 · outbound

This paper cites Tibshirani.An Introduction to the Bootstrap.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Tibshirani.An Introduction to the Bootstrap

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.662854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.689461Z digest=sha256:641121788f33e2ae87c8abc7cdf18fbf6407d81f9fa02d25e377f36cd8d02195

Observation 503b32ec-c227-4ad7-83b1-328a0819f599 · outbound

This paper cites A hybrid model for multilingual ocr.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval A hybrid model for multilingual ocr

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.694040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.694040Z digest=sha256:fa94d31ea07adb63f1141f3bd6b8a15874d59269f24e309879cc06d449d885c4

Observation 03718175-bbd6-463d-ac4a-df5eb644d7bc · outbound

This paper cites Colpali: Efficient document retrieval with vision language models.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Colpali: Efficient document retrieval with vision language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.652851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.698596Z digest=sha256:4dac360c896a426b2f953a170d0dd06cbb15490009be65a23a3d9d9a9d91a874

Observation a61972d4-f41b-4aed-9ad7-4321c77f102a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.703431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.703431Z digest=sha256:a1d21a31588625d3e939a829b3cc22dfa66e463bbff37a8b5b833e7e4edce47e

Observation 16671a50-c346-45eb-bcab-944ba321d299 · outbound

This paper cites Imagebind: One embedding space to bind them all.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Imagebind: One embedding space to bind them all

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.708822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.708822Z digest=sha256:1d964405ad4c70ecf7edc6dca637a670837cacefbca4591924e9e48e4d02394c

Observation 507fc0cc-77dc-41f6-819b-9aa209ddced5 · outbound

This paper cites Cumulated gain-based evaluation of ir techniques.ACM Trans.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Cumulated gain-based evaluation of ir techniques.ACM Trans

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.713608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.713608Z digest=sha256:6c71f0d97eb9f5cf6a4a4052a1c4132f41c37a865038795ab9be3e214c514af7

Observation 552d9585-d1f9-4858-9c82-fde725678c36 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Scaling up visual and vision-language representation learning with noisy text supervision

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.635323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.718360Z digest=sha256:fa456e07de325d47cf4357c169c102cb0dfc9ac0cc3a724750a40ab6f51b4be3

Observation 2fb790ac-e2f2-40a2-bf1c-aebbb741db7a · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.723331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.723331Z digest=sha256:616d215c6bbef630f4391daf62c6874a8a0029ac0fc4285fc3532853946978fc

Observation 7ad71b70-7828-4bcd-926a-111f73edcdc4 · outbound

This paper cites Dense passage retrieval for open-domain question answering.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Dense passage retrieval for open-domain question answering

Reference 18

Resolution
malformed identifier
no resolver link, observed 2026-08-07T06:05:31.728907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.728907Z digest=sha256:63871ad2b5495b56532b21f3662db83e79b31c6fa40645f7a87f452e0da4adcf

Observation 5d801cd2-d676-4a75-ba06-a2f17672d1ac · outbound

This paper cites Colbert: Efficient and effective passage search via con- textualized late interaction over bert.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Colbert: Efficient and effective passage search via con- textualized late interaction over bert

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.734527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.734527Z digest=sha256:265566ca884575eb2afafc927aaca77c47e5005b0feb61417b3e29390bfb68bf

Observation 65563c93-0fa7-45b7-bbe9-24eb6fa33d41 · outbound

This paper cites MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.738189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.738189Z digest=sha256:352fb816fbd35bf454f746820a3d3ff056456b125d13a5a4c5765ef641222075

Observation 7ba335fd-29dc-4963-95c2-4718223f64d4 · outbound

This paper cites Learning to rank for information retrieval.Found.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Learning to rank for information retrieval.Found

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.742958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.742958Z digest=sha256:437f7f05c22d2652b0638b5ae012f5ae0b97d6ec83f7fc0bdeb1d52c04399254

Observation 4a36bfaa-22be-4fb5-92c1-7335d9074463 · outbound

This paper cites CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.747047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.747047Z digest=sha256:8f4c8815f08c5c4d0a770a63ef23348b94bc8d6b153f288a0ae689228c3d87ec

Observation 3a544e91-f00c-4ccf-963a-beedaf9bccc3 · outbound

This paper cites Handwritten optical character recognition (ocr): A comprehensive systematic literature review (slr).IEEE Access, 8: 142642–142668, 2020.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Handwritten optical character recognition (ocr): A comprehensive systematic literature review (slr).IEEE Access, 8: 142642–142668, 2020

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.751704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.751704Z digest=sha256:529a46f5572f408b891edabe842e7057be1a5f7a33292667d357f2b520328429

Observation 3363841f-0da7-4353-aa11-8fee73981574 · outbound

This paper cites Vladva: Discriminative fine-tuning of lvlms,.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Vladva: Discriminative fine-tuning of lvlms,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.624648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.755787Z digest=sha256:0a23b0e1d2b60388af49111dc7017530b0140ca5cbf6cf691dd438afc22f22c4

Observation 58b2133a-98ee-49e3-921d-dad0b7892035 · outbound

This paper cites Learning transferable visual models from natural language supervision.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Learning transferable visual models from natural language supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.764683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.764683Z digest=sha256:4141a911b6ef3cfdfaee40ae2f7262b6246fec8b51ce2694753c6cf71d18b0f7

Observation 642bb0b5-4a42-48eb-aa70-618792b5bc56 · outbound

This paper cites VladVA: Discriminative Fine-tuning of LVLMs.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval VladVA: Discriminative Fine-tuning of LVLMs

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:05:32.158935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.760413Z digest=sha256:103fc061344876c9e85be6e9e75e4434a167cbfc4db74977866325e94b973aaa

Observation 5c79f5b3-dc2d-496f-8114-a8704e143b80 · outbound

This paper cites Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.776301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.776301Z digest=sha256:b0e0b004d8273e26cc69ea125f422ec964ab840e80c171e8747bedb4c46e24db

Observation 5c85e58a-1f3e-4e94-a994-50e76e1fc131 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.780127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.780127Z digest=sha256:c6dbc9ff7b6ef09a6f40bf2adcafde2f53088d6b5181596c25c4e70f34967ae4

Observation 63cd450c-52ea-467c-b862-0b91c26bfd4f · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Robust Speech Recognition via Large-Scale Weak Supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.772749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.772749Z digest=sha256:273808c6433359e6e21c87da9dd4d6c7ba4b2e5d84cdbcad6ff9876daa1b6009

Observation 1ae11b3e-74b7-4d9c-a018-5a43742991c4 · outbound

This paper cites ColBERTv2: Effective and efficient retrieval via lightweight late interaction.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval ColBERTv2: Effective and efficient retrieval via lightweight late interaction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.790322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.790322Z digest=sha256:f984a4411501b87ff80592d8e3f193f2f28d3c467c375103bc9877539bdf66e1

Observation d37da703-0756-46f2-b5ff-e95065c779c1 · outbound

This paper cites an unresolved cited work.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.795910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.795910Z digest=sha256:5773f292e05c7e2eb860fdc01945732a254bf120d2f3df8b641765d2b4c3df04

Observation 2a0a17ba-f607-48f0-887f-96847c25dff2 · outbound

This paper cites MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.784566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.784566Z digest=sha256:eaafe9e620f7ab126fb2a3834c0d6a5814157dce94a7d397dce8401337f1b5b8

Observation 3036bb25-e930-48d7-bee1-519b5b577769 · outbound

This paper cites Gemma 3 Technical Report.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Gemma 3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.804178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.804178Z digest=sha256:5d4143520bac476a709453aab744e52dfdf4337cf177e22afae94acd7410b07d

Observation 326d5b02-42de-418b-bea9-d24b4bf7571a · outbound

This paper cites BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.811569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.811569Z digest=sha256:cd3955e60890bbe7d5e5723b75f7b7c5a9ee761e16e709097573d08e3a132507

Observation 4bcb1f75-7ef4-4d63-beb6-aa3633bf30c1 · outbound

This paper cites Generative multimodal models are in-context learners.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Generative multimodal models are in-context learners

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.601741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.800337Z digest=sha256:1017afad7cb57eb4aabba186fdf1eee767e56abba91807cf7365ab3532123e8f

Observation 77fa6841-8759-45b8-aff8-227fbedaaf7e · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.586001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.819442Z digest=sha256:5c35589449ee7337de0e86a1ef23a3dec216f4dd2e7c0fa24685c3b7a1c414d5

Observation bae2896b-f6e0-4548-b8bc-34a2e87c3ace · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.825064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.825064Z digest=sha256:f753161603497528560b6e0eacd14ee194688bbe6a2de474fa56036c189b25e9

Observation 6f2d055e-5f99-4982-aa01-344b2833129c · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Representation Learning with Contrastive Predictive Coding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.815340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.815340Z digest=sha256:3ce1aefcc2bbaccfdac11df4f34591cb34927cd2f3b7eb15c723bc8b88c0a798

Observation 88bb1e5a-a138-4970-8b01-4361cc5500b2 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.569152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.833276Z digest=sha256:d638a504035f51d39aa8b6f50f5918c736061c08398bb343b206154d96dc7975

Observation 868f1be1-500e-496a-b354-79f63cfaeca3 · outbound

This paper cites Qwen2.5-Omni Technical Report.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Qwen2.5-Omni Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.843978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.843978Z digest=sha256:44fb2ffdb4a0009577c9af0553b971d3174fab937f080891f2b1eeedadfc51cc

Observation d7dda32d-cafc-4d71-9d8f-2e263aedc925 · outbound

This paper cites an unresolved cited work.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.829176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.829176Z digest=sha256:58daad3847591b6740492986c57178875fb582df16dbb8b9354b1463d59b2b68

Observation 91c94b0e-fd37-4767-bc11-35a9a63c640d · outbound

This paper cites UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.852256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.852256Z digest=sha256:1186b579e8b923aaa23806e92c8bdd42c9e9bd3722203428fce50bb8f073eed9

Observation 2677d8dd-e083-4e7c-9b36-385150f719b7 · outbound

This paper cites CREAM: Coarse-to-fine retrieval and multi- modal efficient tuning for document VQA.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval CREAM: Coarse-to-fine retrieval and multi- modal efficient tuning for document VQA

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.537696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.856797Z digest=sha256:ecc637d3fc5697239c92a7c1332e5922a8cc14d826ab960e4599aa649cd4563c

Observation cf2fcc13-053f-43f6-a2b5-69eea1c8d835 · outbound

This paper cites Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment,.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.526800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.861838Z digest=sha256:63b9e6f1f5206309c27d601466a9f42d9a47e663fccf16510a5a51d8197bd5ba

Observation f2675d97-12d0-44d6-b24c-22b115211bfc · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Msr-vtt: A large video description dataset for bridging video and language

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.548489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.847733Z digest=sha256:b878e88d9e751b77c45a55b257ac3fbaca44dff6eadafb1741ba160c492a8df4

Observation 42fd1618-52b1-4a1d-8bde-291cbfa58eaf · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.866869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.866869Z digest=sha256:70978d9530e7a01e61ab40148bd391b6c534c86bf5538ba09cf35f7fdab3a1eb

Observation 79d1128f-70a1-45c3-bf22-9e356735e2a4 · outbound

This paper cites an unresolved cited work.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:32.722824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.648707Z digest=sha256:f4a27eb39497e26de315dc917a63443d22d1fc5d96b12c2e8513dccd4d998e56

Observation 245a8727-c197-4289-ba3c-9bed0af1a4b6 · outbound

This paper cites an unresolved cited work.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:32.558569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:05:31.839371Z digest=sha256:20b76abc5e74f894c01b17f972531c46f339be5c93fb756597d13f02d957413b

Observation 739bd388-8da6-4992-be2b-18cddfdd5aac · outbound

This paper cites Qwen2.5-VL Technical Report.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Qwen2.5-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.640938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.640938Z digest=sha256:5bd340e022605a52c14472ba7301a243534ce020421f8dc7dd9d3ccbff269920

Pith citing papers

Observation 1c39de2c-01db-44ed-a743-64a7eb788bcd · inbound

Hidden in the Multiplicative Interaction: Uncovering Fragility in Multimodal Contrastive Learning cites this paper.

Hidden in the Multiplicative Interaction: Uncovering Fragility in Multimodal Contrastive Learning CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:20:53.354526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:12:07.548447Z digest=sha256:bcc4629791dfd21d945327f31c9da5817ba0165b07759bbd3e8bb258fd6ab8ac