Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T06:05:31.866869Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2506.06144.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T06:05:31.866869Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T19:12:07.548447Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-10T23:20:53.347635Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 19dd6684-388b-4cb7-be7d-caf1c71fb844 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Qwen2.5-vl technical report,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c47f90b-1e70-4dbe-9359-ec944a84271d · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Vlmo: Unified vision- language pre-training with mixture-of-modality-experts
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4d14f39b-27d1-4e85-af09-9a0b17dcb282 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Valor: Vision-audio-language omni-perception pretraining model and dataset, 2023
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 204d3d2f-313a-435e-bdbf-f9d933524ab9 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval V AST: A vision-audio-subtitle-text omni-modality foundation model and dataset
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e6e1dea-47dc-4f43-9209-8bdd96927bd8 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval PaLI: A jointly-scaled multilingual language-image model
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4041aeb2-69d3-403c-9bbd-41e128cb652d · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval M3docrag: Multi- modal retrieval is what you need for multi-page multi-document understanding, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c196a0ea-c869-4dbc-8144-f7bd9a7b07ff · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Jacolbertv2.5: Optimising multi-vector retrievers to create state-of-the-art japanese retrievers with constrained resources, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 823aed3b-6806-41b2-87b7-e11f5308cda9 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Cormack, Charles L A Clarke, and Stefan Buettcher
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdb88050-ddff-42c2-a84a-9587fc3bd8b6 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval QLoRA: Efficient Finetuning of Quantized LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e94752d0-2164-4e0b-b8ee-8b410a9fdcc2 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Tibshirani.An Introduction to the Bootstrap
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 503b32ec-c227-4ad7-83b1-328a0819f599 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval A hybrid model for multilingual ocr
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03718175-bbd6-463d-ac4a-df5eb644d7bc · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Colpali: Efficient document retrieval with vision language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a61972d4-f41b-4aed-9ad7-4321c77f102a · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16671a50-c346-45eb-bcab-944ba321d299 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Imagebind: One embedding space to bind them all
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 507fc0cc-77dc-41f6-819b-9aa209ddced5 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Cumulated gain-based evaluation of ir techniques.ACM Trans
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 552d9585-d1f9-4858-9c82-fde725678c36 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Scaling up visual and vision-language representation learning with noisy text supervision
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2fb790ac-e2f2-40a2-bf1c-aebbb741db7a · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ad71b70-7828-4bcd-926a-111f73edcdc4 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Dense passage retrieval for open-domain question answering
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d801cd2-d676-4a75-ba06-a2f17672d1ac · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Colbert: Efficient and effective passage search via con- textualized late interaction over bert
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65563c93-0fa7-45b7-bbe9-24eb6fa33d41 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ba335fd-29dc-4963-95c2-4718223f64d4 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Learning to rank for information retrieval.Found
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a36bfaa-22be-4fb5-92c1-7335d9074463 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a544e91-f00c-4ccf-963a-beedaf9bccc3 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Handwritten optical character recognition (ocr): A comprehensive systematic literature review (slr).IEEE Access, 8: 142642–142668, 2020
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3363841f-0da7-4353-aa11-8fee73981574 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Vladva: Discriminative fine-tuning of lvlms,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 58b2133a-98ee-49e3-921d-dad0b7892035 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Learning transferable visual models from natural language supervision
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 642bb0b5-4a42-48eb-aa70-618792b5bc56 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval VladVA: Discriminative Fine-tuning of LVLMs
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c79f5b3-dc2d-496f-8114-a8704e143b80 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c85e58a-1f3e-4e94-a994-50e76e1fc131 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63cd450c-52ea-467c-b862-0b91c26bfd4f · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Robust Speech Recognition via Large-Scale Weak Supervision
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ae11b3e-74b7-4d9c-a018-5a43742991c4 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval ColBERTv2: Effective and efficient retrieval via lightweight late interaction
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d37da703-0756-46f2-b5ff-e95065c779c1 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a0a17ba-f607-48f0-887f-96847c25dff2 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3036bb25-e930-48d7-bee1-519b5b577769 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Gemma 3 Technical Report
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 326d5b02-42de-418b-bea9-d24b4bf7571a · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bcb1f75-7ef4-4d63-beb6-aa3633bf30c1 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Generative multimodal models are in-context learners
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 77fa6841-8759-45b8-aff8-227fbedaaf7e · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bae2896b-f6e0-4548-b8bc-34a2e87c3ace · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f2d055e-5f99-4982-aa01-344b2833129c · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Representation Learning with Contrastive Predictive Coding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88bb1e5a-a138-4970-8b01-4361cc5500b2 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Longvideobench: A benchmark for long-context interleaved video-language understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 868f1be1-500e-496a-b354-79f63cfaeca3 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Qwen2.5-Omni Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7dda32d-cafc-4d71-9d8f-2e263aedc925 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91c94b0e-fd37-4767-bc11-35a9a63c640d · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2677d8dd-e083-4e7c-9b36-385150f719b7 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval CREAM: Coarse-to-fine retrieval and multi- modal efficient tuning for document VQA
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf2fcc13-053f-43f6-a2b5-69eea1c8d835 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2675d97-12d0-44d6-b24c-22b115211bfc · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Msr-vtt: A large video description dataset for bridging video and language
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42fd1618-52b1-4a1d-8bde-291cbfa58eaf · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79d1128f-70a1-45c3-bf22-9e356735e2a4 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Unresolved cited work
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 245a8727-c197-4289-ba3c-9bed0af1a4b6 · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 739bd388-8da6-4992-be2b-18cddfdd5aac · outbound
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Qwen2.5-VL Technical Report
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c39de2c-01db-44ed-a743-64a7eb788bcd · inbound
Hidden in the Multiplicative Interaction: Uncovering Fragility in Multimodal Contrastive Learning CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.