Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:46.090733Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 100 of 120 outbound references and 12 inbound Pith citation observations for arXiv:2505.19650.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:46.090733Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T18:30:14.559722Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T05:49:36.602258Z
100 of 120 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d25bdf05-a782-4319-83d6-0091c3fc8af7 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Multimodal automated fact-checking: A survey
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62a061c5-05f1-48aa-9003-4f4ff0a5f489 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Localizing moments in video with natural language
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1880f482-8f26-4197-a149-8fc3aa7ba92d · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Vqa: Visual question answering
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 613f3f9a-f2c5-4326-be34-a86460cd5d22 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47fa31f2-a075-4f07-b00d-96c42d4efb7c · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Qwen2.5-VL Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b81a194e-645b-4e3f-a99d-899caf53458c · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33c681ea-322c-4d07-9570-bf88a7f1e689 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Wiki-llava: Hierarchical retrieval-augmented generation for multimodal llms
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caca77dd-0c7f-436f-a4a7-f8dbfe66e7db · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e2aef62-91dd-4624-9d83-0869ab6c18e7 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Webqa: Multihop and multimodal qa
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a17a1232-f62d-4dfb-8d2d-06f2228c2af9 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Collecting highly parallel data for paraphrase evaluation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14038995-ac3d-40b0-9f9d-34599cad0315 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ce443e8-645b-4c1c-864b-3fc15153c6aa · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Sharegpt4v: Improving large multi-modal models with better captions
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb626853-7142-459a-bb3f-07ee6af82e55 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3abfb845-1d29-48d6-909d-845825214698 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ff154db-8bbf-46bf-91b2-fba69349d5dd · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55ade486-db11-43f2-b8c3-e76b36b8c472 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Reproducible scaling laws for contrastive language-image learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce0df1fa-2f49-48a6-be40-6c7c99cc4813 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Visual dialog
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e03ce079-de1e-42e5-bb72-41a602d79443 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Imagenet: A large-scale hierarchical image database
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20feab99-8379-442e-816a-dc2af501376d · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Mmdocir: Benchmarking multi-modal retrieval for long documents.arXiv preprint arXiv:2501.08828,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b9a2408-e4ac-48f3-913d-a09e84c5684c · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval The pascal visual object classes challenge: A retrospective.International Journal of Computer Vision, 111:98–136, 2015
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a452dfd7-f6d5-419d-b6bb-18c1ba7ec49e · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56c6cc8f-ded3-4f71-a5b4-ad3cb8f4bb2b · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Simcse: Simple contrastive learning of sentence embeddings
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f0689e4-5de1-42e4-b919-de06d58519da · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Imagebind: One embedding space to bind them all
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ada8643-244a-4f40-9648-261e9c3c503a · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Breaking the modality barrier: Universal embedding learning with multimodal llms.arXiv preprint arXiv:2504.17432, 2025
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e76bd2ca-1ea7-42db-a61f-afaabef2ee8c · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Lora: Low-rank adaptation of large language models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d97d084b-ec3b-48e1-9ac2-4f3a6cb753ed · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Egocvr: An egocentric benchmark for fine-grained composed video retrieval
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf2eed1-d342-476f-b3c2-12fcc2223410 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Mate: Meet at the embedding-connecting images with long texts
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1932084e-9cc0-4812-8e24-96947c515a18 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Scaling up visual and vision-language representation learning with noisy text supervision
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d567483b-84af-428a-9f33-3baca666fa76 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Tencent text- video retrieval: hierarchical cross-modal interactions with multi-level representations.IEEE Access, 2022
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69ea240f-4921-45fd-bd9a-8d749f023325 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Scaling sentence embeddings with large language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c11c69d-5b64-4c24-bd72-49451fd8d076 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval E5-V: Universal Embeddings with Multimodal Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac76a1ed-3331-45f8-803f-635eead6ea98 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67ea4878-8040-4681-b778-07b3235699d1 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92d103e5-e08d-4319-b296-89169c5ffde4 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Visual question answering: Datasets, algorithms, and future challenges.Computer Vision and Image Understanding, 163:3–20, 2017
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 291c51a4-7d5c-46bf-b507-49027c432b08 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Deep visual-semantic alignments for generating image descriptions
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0811a0e4-3d0c-4eb8-be38-e8d78780461e · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval The hateful memes challenge: Detecting hate speech in multimodal memes
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0a6ae9e-06f8-4ccb-b4df-16bb641e9b7d · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e1e8e44-c342-4aad-a0f8-f9dd659c04b4 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466, 2019
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28885696-1d45-4d01-9b13-455c7adb32cc · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Llave: Large language and vision embedding models with hardness-weighted contrastive learning.arXiv preprint arXiv:2503.04812, 2025
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 655aff41-ce43-41c3-8937-eba3fd1e8898 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval LLaVA-OneVision: Easy Visual Task Transfer
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59753726-70b8-44bb-8bef-ba75c75e89e8 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a365334-6104-47de-ad03-3b4afc2900d8 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f723ea6-4cd0-4cd6-a9d8-b15993af6316 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6817cad-c2f1-499f-afc9-97037fc12468 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67f62fc2-a81e-49fa-a515-6d3472ac5dd0 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Microsoft coco: Common objects in context
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 551991ec-7975-4641-a0f8-c7238f90f6df · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e83cd7ee-0e9e-4b0d-a869-07c9dd7e76c4 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Visual News: Benchmark and Challenges in News Image Captioning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c57bc6-8e15-40d9-96aa-0c7f592a4b37 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Improved baselines with visual instruction tuning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 659449bb-740b-47ae-9b37-96afcaccccd4 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f85b993-5da8-47b3-b77c-2a1152f5dbe7 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Visual instruction tuning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7908fffa-9e17-4bd0-af99-c05ba6b87985 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Llava-plus: Learning to use tools for creating multimodal agents
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 923723cb-467f-45db-a2a1-8559ab4dc228 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 260b0014-20bc-4880-bd2e-ddd6d1537b45 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Image retrieval on real-life images with pre-trained vision-and-language models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a94f9504-f8e7-4bac-8e61-eff261c60a71 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Generative multi-modal knowledge retrieval with large language models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee3f2b50-ba0a-43d5-a43c-54a563afc47a · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval End-to-end knowl- edge retrieval with multi-modal queries
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eeb2464-83a1-42c9-a2e2-7903dc66dde8 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Video-rag: Visually-aligned retrieval-augmented long video comprehension.arXiv preprint arXiv:2411.13093, 2024
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b2f095a-1e1e-4972-9c8b-953e279d9a02 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Unifying multimodal retrieval via document screenshot embedding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8e194b3-da69-47a2-a2ad-e01cec6593b6 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ab821c6-af25-475e-8c26-7541f1f79ee4 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2b4ea5b-3fde-419d-a747-7179956369f7 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 023bd09b-026c-499f-8603-90db05775295 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Infographicvqa
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c140d83b-3342-4e5b-a369-99b7ede44866 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Docvqa: A dataset for vqa on document images
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d78713e-fe65-4b1c-b9c1-1e0f20828970 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Mm1: methods, analysis and insights from multimodal llm pre-training
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ced10b56-ccc6-4d8e-a033-eb1cce907549 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Docci: Descriptions of connected and contrasting images
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 49df6ac7-3954-48eb-b100-f5fc6a1d8b9b · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Representation Learning with Contrastive Predictive Coding
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b9d012b-7ae5-462b-905a-907bb39c1e1c · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Kosmos-G: Generating Images in Context with Multimodal Large Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4c51ae5-fb47-48a0-91db-1ed9642c553b · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5496358f-8197-43d5-854a-cab5644921b9 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Filtering, distillation, and hard negatives for vision-language pre-training
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ca105a17-df5a-4e29-84df-50b64bd82dc9 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Learning transferable visual models from natural language supervision
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e597097e-88f1-46d8-b812-e0d643ff95d8 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Squad: 100,000+ questions for machine comprehension of text
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e021a51d-e594-4c22-9eb9-a07b0d2b105e · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Contrastive Learning with Hard Negative Samples
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 465889c8-cf34-48ac-9603-90480d251467 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Pic2word: Mapping pictures to words for zero-shot composed image retrieval
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9c0da33c-6771-4f97-be58-70b508e1a057 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Laion- 5b: An open large-scale dataset for training next generation image-text models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 410094f7-2180-4340-a721-0811d9523df3 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval A-okvqa: A benchmark for visual question answering using world knowledge
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1393e73d-2486-4a8e-9db2-783e9b418355 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fc5fba1-6b55-444f-83ac-485e825a272f · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Composed video retrieval via enriched context and discriminative embeddings
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9ddbad45-f0bd-4905-afec-10a0b6c37f36 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Fever: a large-scale dataset for fact extraction and verification
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7e49834f-8235-49d4-ab1a-589f9b41a9d2 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Covr-2: Automatic data construction for composed video retrieval.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 43ac363e-3f70-45b1-b50d-cb55d8b92e58 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Covr: Learning composed video retrieval from web video captions
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4165373e-c41e-4f73-bb44-aa18251c6109 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa90b388-4267-4933-acb7-be02d1deee11 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval A Comprehensive Survey on Cross-modal Retrieval
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 112d2fbc-dfce-415c-9cee-36e5875b4c65 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00e1ee86-7f82-4c3e-9637-2ac103fc5528 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Fvqa: Fact-based visual question answering.IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(10):2413–2427, 2017
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f7e16bdf-8b49-4887-aea7-5fe6a581e0f8 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Cross- modal retrieval: a systematic review of methods and future directions.Proceedings of the IEEE, 2025
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 22241213-305e-4750-a912-386dc00fafc9 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Internvid: A large-scale video-text dataset for multimodal understanding and generation
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1ccdfecc-be1c-4081-870d-5ebf432c8179 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Internvideo2: Scaling foundation models for multimodal video understanding
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c92fe79b-13d1-4bc5-a59a-c97df7637a2d · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f54c8e76-bdc0-480b-9fc8-a497a0ec08de · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e20025dc-ed00-4245-aed1-d7b7f54c7576 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dc596e4-9b13-4caa-9b6a-29a932e200f5 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval N24news: A new dataset for multimodal news classification
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2c6dc2a6-8756-4a3a-bc80-de2df8aaa6c1 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Uniir: Training and benchmarking universal multimodal information retrievers
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0ccbf610-017b-47cf-a0cf-93ed9d61fcfd · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Sun database: Large-scale scene recognition from abbey to zoo
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b10f7c85-ecde-490e-acd6-857b03f829f8 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fafc94c8-f7eb-405c-bbee-bdbf296ecd29 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Msr-vtt: A large video description dataset for bridging video and language
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 731ad417-6323-44dd-8cd0-da784f786d16 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff5b82ca-a31b-43ae-aebe-5614bf80505d · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f71f51c-5bf3-4289-8e15-5670d5225064 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0da0f99f-033f-4503-b951-6c6d25f52ded · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval End-to-end multimodal fact-checking and explanation generation: A challenging dataset and models
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0b5c142e-69e9-4709-91ce-7bed037176bb · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25d936f7-e057-4d04-b373-1597cff5eec4 · outbound
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee1ebd9e-a9c9-45ae-ae6a-38cd48861805 · inbound
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edbdef5e-892a-46ba-bff0-03e3b5d5177a · inbound
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 913342ff-b8a1-4a73-8d70-b9dd8edb213b · inbound
FreeRet: MLLMs as Training-Free Retrievers Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 53c0bd72-bc9e-42d9-87d1-38ee6fc94016 · inbound
FreeRet: MLLMs as Training-Free Retrievers Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa897247-7e0b-4148-9f23-44169316f1bc · inbound
Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8f2ce394-f8b1-45e9-893c-e0de3bc7989f · inbound
Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c769451-754c-4c00-a1c5-adf5685cc380 · inbound
Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a9ccc3a-2e9a-4e11-9033-1cd3626b5ee9 · inbound
Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8a5273d6-52b2-4e9e-9818-5a94e51accd8 · inbound
SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19bd5014-640a-4ada-a4d1-8c375f21155c · inbound
ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 68e4c1a0-4d0d-436a-b85c-8bffd421c7ac · inbound
Illuminating Visual Identity in Universal Multimodal Embeddings Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ba9a618-5492-42cf-be31-7c0106c94e0c · inbound
Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.