Pith. sign in

Paper Citation Record · LEDGER

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

As of 8 August 2026, this Paper Citation Record lists 100 of 120 outbound references and 12 inbound Pith citation observations for arXiv:2505.19650.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19650 v2

Coverage vector

measured 100 of 120 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:46.090733Z

measured 112 of 112 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:30:14.559722Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T05:49:36.602258Z

Reference resolution

100 of 120 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved78
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d25bdf05-a782-4319-83d6-0091c3fc8af7 · outbound

This paper cites Multimodal automated fact-checking: A survey.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Multimodal automated fact-checking: A survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.135801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.135801Z digest=sha256:7620ba4567799789f8d16986e40911dbe3f56aee411502b5c4ef1fb67c2fcabc

Observation 62a061c5-05f1-48aa-9003-4f4ff0a5f489 · outbound

This paper cites Localizing moments in video with natural language.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Localizing moments in video with natural language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.194753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.194753Z digest=sha256:4bf310ad455b6c403148e949b3ab59a944c9ff170f4d280e4433d0d70569f2bf

Observation 1880f482-8f26-4197-a149-8fc3aa7ba92d · outbound

This paper cites Vqa: Visual question answering.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Vqa: Visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.236558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.236558Z digest=sha256:d0a59bdc4b6aab6c87df4ac911c94cc579ab1167ea9ec0221fd63db6f331804f

Observation 613f3f9a-f2c5-4326-be34-a86460cd5d22 · outbound

This paper cites FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.333584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.333584Z digest=sha256:84ec07416f0b741d8931f9c927f7665acf99ce36fba3b48930287b80e0734b44

Observation 47fa31f2-a075-4f07-b00d-96c42d4efb7c · outbound

This paper cites Qwen2.5-VL Technical Report.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.445117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.445117Z digest=sha256:0aab1f8e21fcf4e9f5ca7b5ab59dfd425b8c9a153e5798ad3ce7cd6f0dd16234

Observation b81a194e-645b-4e3f-a99d-899caf53458c · outbound

This paper cites MS MARCO: A Human Generated MAchine Reading COmprehension Dataset.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval MS MARCO: A Human Generated MAchine Reading COmprehension Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.507574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.507574Z digest=sha256:e645f79ad781119611790ecbbd156a60e49003ff6a777d40bec0d0ae9ab0e3b4

Observation 33c681ea-322c-4d07-9570-bf88a7f1e689 · outbound

This paper cites Wiki-llava: Hierarchical retrieval-augmented generation for multimodal llms.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Wiki-llava: Hierarchical retrieval-augmented generation for multimodal llms

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.583688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.583688Z digest=sha256:ea189cce69cd3ec73f48a247d49510cd4f81c7a88665d612aab0fa118f476f86

Observation caca77dd-0c7f-436f-a4a7-f8dbfe66e7db · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.686789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.686789Z digest=sha256:c65103cec573fe019cc7327d535083a845d2f87fdfb5fd80a15ee96eaec719bd

Observation 8e2aef62-91dd-4624-9d83-0869ab6c18e7 · outbound

This paper cites Webqa: Multihop and multimodal qa.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Webqa: Multihop and multimodal qa

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.792978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.792978Z digest=sha256:66481004cdb8d38a44313066436819ca17f468a8ab61497d34092897277f988e

Observation a17a1232-f62d-4dfb-8d2d-06f2228c2af9 · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Collecting highly parallel data for paraphrase evaluation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.877735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.877735Z digest=sha256:9e4b0a043371bad1ff213ed1970613181fdaa3663a6c3cd9d7fd05848215bee4

Observation 14038995-ac3d-40b0-9f9d-34599cad0315 · outbound

This paper cites mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.978957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.978957Z digest=sha256:e60f0c8815d830e71d0e99452a770c6926818f467291b9a344a13db0eef7b093

Observation 5ce443e8-645b-4c1c-864b-3fc15153c6aa · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Sharegpt4v: Improving large multi-modal models with better captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.039664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.039664Z digest=sha256:98f62767137c06dfa878608d245935c6d66dd6dc72442305cbeed64f61fe6213

Observation cb626853-7142-459a-bb3f-07ee6af82e55 · outbound

This paper cites Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.135154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.135154Z digest=sha256:ec5ca771faff3aafb38106106ec37f8e5e1010298f3127cc369996c41c0a080f

Observation 3abfb845-1d29-48d6-909d-845825214698 · outbound

This paper cites MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.241872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.241872Z digest=sha256:21000d8b06efc61771de134b49b1126f61bfcf63bf1f70bb06ad26cb5ab7ad67

Observation 2ff154db-8bbf-46bf-91b2-fba69349d5dd · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.335797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.335797Z digest=sha256:13513c38d295402159ec9164d7b4ba8107fcf6bce3eb5f4fa036628f35d19af9

Observation 55ade486-db11-43f2-b8c3-e76b36b8c472 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Reproducible scaling laws for contrastive language-image learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.353327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.353327Z digest=sha256:12683933fda7e362c74f9e674e458dfd177fbdcc50423f42892d44f90ef59e87

Observation ce0df1fa-2f49-48a6-be40-6c7c99cc4813 · outbound

This paper cites Visual dialog.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Visual dialog

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.487073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.487073Z digest=sha256:d58d48231df5816a1aeb175d789e1704efad2e47d9638c7e774b57db9355aeb9

Observation e03ce079-de1e-42e5-bb72-41a602d79443 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Imagenet: A large-scale hierarchical image database

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.611355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.611355Z digest=sha256:57c6fd729f3a5e748b1666dab285962c5659ccb7c1319950f14cd88b954b7b98

Observation 20feab99-8379-442e-816a-dc2af501376d · outbound

This paper cites Mmdocir: Benchmarking multi-modal retrieval for long documents.arXiv preprint arXiv:2501.08828,.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Mmdocir: Benchmarking multi-modal retrieval for long documents.arXiv preprint arXiv:2501.08828,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.724894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.724894Z digest=sha256:1d0841402dd269c41e454f5c19bd43de234f4e5830db9a91868ff13b26967040

Observation 6b9a2408-e4ac-48f3-913d-a09e84c5684c · outbound

This paper cites The pascal visual object classes challenge: A retrospective.International Journal of Computer Vision, 111:98–136, 2015.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval The pascal visual object classes challenge: A retrospective.International Journal of Computer Vision, 111:98–136, 2015

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.857313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.857313Z digest=sha256:7b4a2edec1d4d1799e3b5e6ca818332e94b710355effd0a22031df069edea2ae

Observation a452dfd7-f6d5-419d-b6bb-18c1ba7ec49e · outbound

This paper cites DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.993221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.993221Z digest=sha256:339bcd640304d140a71f042a68c4fe9bc6b9e7b6da8f22b5183aa576a168f54b

Observation 56c6cc8f-ded3-4f71-a5b4-ad3cb8f4bb2b · outbound

This paper cites Simcse: Simple contrastive learning of sentence embeddings.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Simcse: Simple contrastive learning of sentence embeddings

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:39.124769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:39.124769Z digest=sha256:5fda0d761444c6bc55c08526c856e9c55f77c07896b438bac202b4d3a1aefe1a

Observation 7f0689e4-5de1-42e4-b919-de06d58519da · outbound

This paper cites Imagebind: One embedding space to bind them all.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Imagebind: One embedding space to bind them all

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:39.289062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:39.289062Z digest=sha256:f59deb99042938fb1f4e5f116d735c2a97b88f7e4e997083791254187b679eb6

Observation 9ada8643-244a-4f40-9648-261e9c3c503a · outbound

This paper cites Breaking the modality barrier: Universal embedding learning with multimodal llms.arXiv preprint arXiv:2504.17432, 2025.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Breaking the modality barrier: Universal embedding learning with multimodal llms.arXiv preprint arXiv:2504.17432, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:39.416668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:39.416668Z digest=sha256:00b2a9292120d4b9bbfeeb666183f9ed2b85e0e795735954fb97160b4d4093ba

Observation e76bd2ca-1ea7-42db-a61f-afaabef2ee8c · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Lora: Low-rank adaptation of large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:39.525902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:39.525902Z digest=sha256:66511d6ad57156cff63c3bd09d05a34b148b0a29d8ed07c70a107a36bbd6c30c

Observation d97d084b-ec3b-48e1-9ac2-4f3a6cb753ed · outbound

This paper cites Egocvr: An egocentric benchmark for fine-grained composed video retrieval.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Egocvr: An egocentric benchmark for fine-grained composed video retrieval

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:39.674376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:39.674376Z digest=sha256:761eadadb1730c2a5a90436189fa82f1e8b0585693541b4f14374bff19f6e021

Observation adf2eed1-d342-476f-b3c2-12fcc2223410 · outbound

This paper cites Mate: Meet at the embedding-connecting images with long texts.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Mate: Meet at the embedding-connecting images with long texts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:39.787447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:39.787447Z digest=sha256:aea9714e7bce37844eab90179f64e6e4e21f29a56bd2a207d8e2d6411b423325

Observation 1932084e-9cc0-4812-8e24-96947c515a18 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Scaling up visual and vision-language representation learning with noisy text supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:39.943808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:39.943808Z digest=sha256:6ebb9e885b09bbe554d1873e94ba149592c0e8806417d1a43f6d3fdc268ce3b7

Observation d567483b-84af-428a-9f33-3baca666fa76 · outbound

This paper cites Tencent text- video retrieval: hierarchical cross-modal interactions with multi-level representations.IEEE Access, 2022.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Tencent text- video retrieval: hierarchical cross-modal interactions with multi-level representations.IEEE Access, 2022

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.007625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.007625Z digest=sha256:38e40d97691945d30dc699d25994c8c182bfe0765dafe0793a1cf5744824f5cf

Observation 69ea240f-4921-45fd-bd9a-8d749f023325 · outbound

This paper cites Scaling sentence embeddings with large language models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Scaling sentence embeddings with large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.062977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.062977Z digest=sha256:040f0aa68adadb76a3ec6f293376bbb586f1b124f420d438a8ce647951ea788f

Observation 7c11c69d-5b64-4c24-bd72-49451fd8d076 · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.135649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.135649Z digest=sha256:aa83410d927a00f9781af38ed80d3418b3699e3661b0ba7ee9f0b9c97235b240

Observation ac76a1ed-3331-45f8-803f-635eead6ea98 · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.233513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.233513Z digest=sha256:c6d1698fecd2500373bd90815b467f6e0864f865d7240e2ff31bf73e37e69ddf

Observation 67ea4878-8040-4681-b778-07b3235699d1 · outbound

This paper cites Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.292912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.292912Z digest=sha256:a80da7320c312f68a140bc98f29788e67368d742600d3aa3037691e30be15349

Observation 92d103e5-e08d-4319-b296-89169c5ffde4 · outbound

This paper cites Visual question answering: Datasets, algorithms, and future challenges.Computer Vision and Image Understanding, 163:3–20, 2017.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Visual question answering: Datasets, algorithms, and future challenges.Computer Vision and Image Understanding, 163:3–20, 2017

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.422494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.422494Z digest=sha256:97c31f5fb71c2c6b348f890fe12dfec450449b505ec708beec56c71c0dcec8c4

Observation 291c51a4-7d5c-46bf-b507-49027c432b08 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Deep visual-semantic alignments for generating image descriptions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.496653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.496653Z digest=sha256:e46fc2e583e18cacb87cffd028d6f15833aa81a48bee4169784761c91371e388

Observation 0811a0e4-3d0c-4eb8-be38-e8d78780461e · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.596642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.596642Z digest=sha256:cc457a60a3e2477ebc74dbb8189319193eed0e424706ab10fad7de1ba4ebacb8

Observation e0a6ae9e-06f8-4ccb-b4df-16bb641e9b7d · outbound

This paper cites TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.649758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.649758Z digest=sha256:0313231628401db45ba4f5bb4d8dedbe76710d28b007b95790532fe44df60419

Observation 6e1e8e44-c342-4aad-a0f8-f9dd659c04b4 · outbound

This paper cites Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466, 2019.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466, 2019

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.710060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.710060Z digest=sha256:903106a2de7f05a8c2b86a24d59f76835204e3c7899e157f8c329095a9cbed3e

Observation 28885696-1d45-4d01-9b13-455c7adb32cc · outbound

This paper cites Llave: Large language and vision embedding models with hardness-weighted contrastive learning.arXiv preprint arXiv:2503.04812, 2025.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Llave: Large language and vision embedding models with hardness-weighted contrastive learning.arXiv preprint arXiv:2503.04812, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.831458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.831458Z digest=sha256:e1c574967ff157a70caf853b7c13d6d6fe721d0ffec5e12c6ae4299e996dfd20

Observation 655aff41-ce43-41c3-8937-eba3fd1e8898 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval LLaVA-OneVision: Easy Visual Task Transfer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.926888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.926888Z digest=sha256:5ab9ba9b3b4a986ab68ef70fb00e595ab4a9b46717dea3e9b59b16f7ce1d1091

Observation 59753726-70b8-44bb-8bef-ba75c75e89e8 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.027258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.027258Z digest=sha256:7bb531906ec16b999894784e8d0ba32c0b7acd33823c1e2e5c69cbe9c8d5e5bc

Observation 5a365334-6104-47de-ad03-3b4afc2900d8 · outbound

This paper cites Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.096391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.096391Z digest=sha256:2cf23774fbfa4a4d113b01c50aed3ccbf5c5e3299adcf8e7e335381d33b9f118

Observation 9f723ea6-4cd0-4cd6-a9d8-b15993af6316 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.208727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.208727Z digest=sha256:494e9ea197ab050fd44d5d32c4c5026befe3748d7967f366dc77346d6eff612d

Observation a6817cad-c2f1-499f-afc9-97037fc12468 · outbound

This paper cites MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.299337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.299337Z digest=sha256:de9196f250b284a297a6bdefbcb9ff7ad98b6cc907e8199c99365a5a9a018308

Observation 67f62fc2-a81e-49fa-a515-6d3472ac5dd0 · outbound

This paper cites Microsoft coco: Common objects in context.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Microsoft coco: Common objects in context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.369793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.369793Z digest=sha256:c1cefb8364582a5eeedd7cfd1c0005f1ef4aca25bdfc0294af8fbe51cdb30aad

Observation 551991ec-7975-4641-a0f8-c7238f90f6df · outbound

This paper cites IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.513558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.513558Z digest=sha256:c7dda6753822b18f223d38a12cc68fabd2a080c34ac6c21138578522ab811093

Observation e83cd7ee-0e9e-4b0d-a869-07c9dd7e76c4 · outbound

This paper cites Visual News: Benchmark and Challenges in News Image Captioning.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Visual News: Benchmark and Challenges in News Image Captioning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.573083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.573083Z digest=sha256:48225afa9c7aca62ab826150e77aeecf232a33f04ff742b5c43e04bf76945922

Observation b3c57bc6-8e15-40d9-96aa-0c7f592a4b37 · outbound

This paper cites Improved baselines with visual instruction tuning.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Improved baselines with visual instruction tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.612473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.612473Z digest=sha256:8dcbfaa8d729d6ed237028d82618f9b53ae8043f6737a07df8163eb02d7277c7

Observation 659449bb-740b-47ae-9b37-96afcaccccd4 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.651069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.651069Z digest=sha256:002678dc00dc5e14f98c1918ba3e847301b3600d0483fcb176c9f3513612e239

Observation 8f85b993-5da8-47b3-b77c-2a1152f5dbe7 · outbound

This paper cites Visual instruction tuning.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Visual instruction tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.752801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.752801Z digest=sha256:9231339a9f2594283a59b931b07b9cc8034a2d1f465f37e39e1a7ca203239b12

Observation 7908fffa-9e17-4bd0-af99-c05ba6b87985 · outbound

This paper cites Llava-plus: Learning to use tools for creating multimodal agents.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Llava-plus: Learning to use tools for creating multimodal agents

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.844748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.844748Z digest=sha256:2177668cd2ccf9053f1e13623eb49397043e3269b6d00fb6c3e9410288dde30d

Observation 923723cb-467f-45db-a2a1-8559ab4dc228 · outbound

This paper cites LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.923753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.923753Z digest=sha256:5c277c67487611fecd81e6605594feb12cc17c00734c17874d1e8d2d3705dfe4

Observation 260b0014-20bc-4880-bd2e-ddd6d1537b45 · outbound

This paper cites Image retrieval on real-life images with pre-trained vision-and-language models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Image retrieval on real-life images with pre-trained vision-and-language models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.986061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.986061Z digest=sha256:9043d3ee914e7cd27b740f79b8f224b469c88a3b4bbe6f269fd2035c95378aba

Observation a94f9504-f8e7-4bac-8e61-eff261c60a71 · outbound

This paper cites Generative multi-modal knowledge retrieval with large language models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Generative multi-modal knowledge retrieval with large language models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.063940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.063940Z digest=sha256:2526869d0efa5588712227470ca1fc48df35efa440e6dc314225f38847b90752

Observation ee3f2b50-ba0a-43d5-a43c-54a563afc47a · outbound

This paper cites End-to-end knowl- edge retrieval with multi-modal queries.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval End-to-end knowl- edge retrieval with multi-modal queries

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.139537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.139537Z digest=sha256:f9021dbe9595b96c4f2605e94147075cdb9a141345ed7b829400c1261969375e

Observation 6eeb2464-83a1-42c9-a2e2-7903dc66dde8 · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension.arXiv preprint arXiv:2411.13093, 2024.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Video-rag: Visually-aligned retrieval-augmented long video comprehension.arXiv preprint arXiv:2411.13093, 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.196435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.196435Z digest=sha256:df9a343d2dbce193ccebf2b3c318e25b2dab2af39199d45b7e7787bea8db4172

Observation 6b2f095a-1e1e-4972-9c8b-953e279d9a02 · outbound

This paper cites Unifying multimodal retrieval via document screenshot embedding.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Unifying multimodal retrieval via document screenshot embedding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.277212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.277212Z digest=sha256:ecbb0e4995b6fc8e117ca9750eccac6c6bfeb2c4711ef404d7ee2aa48596427f

Observation d8e194b3-da69-47a2-a2ad-e01cec6593b6 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.381843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.381843Z digest=sha256:155b5297f0f92e0f1e3cfd47c247c4438d0fbcf847f94010829a1ebb4594f12f

Observation 4ab821c6-af25-475e-8c26-7541f1f79ee4 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.469662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.469662Z digest=sha256:a9a899c11487756a665f2ccd7ec7c7b7face6e37ac39eac02b2913f30011bebc

Observation e2b4ea5b-3fde-419d-a747-7179956369f7 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.534731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.534731Z digest=sha256:24f4267c819cf5e182c1608cb46d17055212db02a6a4143766abb581b9b1fe3a

Observation 023bd09b-026c-499f-8603-90db05775295 · outbound

This paper cites Infographicvqa.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Infographicvqa

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.612365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.612365Z digest=sha256:90c40e452959e96228c60c0f2a2a856923b0fb0b0dbc8fa48ed3087cad6c10d8

Observation c140d83b-3342-4e5b-a369-99b7ede44866 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Docvqa: A dataset for vqa on document images

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.664676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.664676Z digest=sha256:a030c389f9070ed3f270b50aaad136c892dbba992a934bef494ecbdf22005eb2

Observation 4d78713e-fe65-4b1c-b9c1-1e0f20828970 · outbound

This paper cites Mm1: methods, analysis and insights from multimodal llm pre-training.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Mm1: methods, analysis and insights from multimodal llm pre-training

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:55.493208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:42.732833Z digest=sha256:d385bb94258d0591f0b1853ae60aa846ba720481fe24aaa94d857b9036e26b5d

Observation ced10b56-ccc6-4d8e-a033-eb1cce907549 · outbound

This paper cites Docci: Descriptions of connected and contrasting images.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Docci: Descriptions of connected and contrasting images

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:55.391113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:42.803646Z digest=sha256:3430bf38ddb6ccb64940555ca8f7556dd3c57aca27f0e7d273a3dd62a4f6a90e

Observation 49df6ac7-3954-48eb-b100-f5fc6a1d8b9b · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Representation Learning with Contrastive Predictive Coding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.931680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.931680Z digest=sha256:9ddc26b1a90c597f8bf182d5a0a6954f20273a7e6b0722b45d9b119145f947a7

Observation 0b9d012b-7ae5-462b-905a-907bb39c1e1c · outbound

This paper cites Kosmos-G: Generating Images in Context with Multimodal Large Language Models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.998277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.998277Z digest=sha256:7dae42820bf09cd031957108654e03f8a015332564f6028de9fc29709c8cd3ba

Observation d4c51ae5-fb47-48a0-91db-1ed9642c553b · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:55.208547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:43.085408Z digest=sha256:332c46b9d3c36ea492a5ae24266871362aeb92fd27f9c07e3bcbace152f31909

Observation 5496358f-8197-43d5-854a-cab5644921b9 · outbound

This paper cites Filtering, distillation, and hard negatives for vision-language pre-training.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Filtering, distillation, and hard negatives for vision-language pre-training

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:55.069133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:43.158713Z digest=sha256:c84070391f26024c3cd4bf9755f075faa8bc98071983d1772aaea1dd4fcc021f

Observation ca105a17-df5a-4e29-84df-50b64bd82dc9 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Learning transferable visual models from natural language supervision

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:54.824769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:43.236786Z digest=sha256:d3cb3812dc6e6fc611d330881e71f09eac1141babdcd7c12eed5892040888d33

Observation e597097e-88f1-46d8-b812-e0d643ff95d8 · outbound

This paper cites Squad: 100,000+ questions for machine comprehension of text.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Squad: 100,000+ questions for machine comprehension of text

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:54.517715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:43.286893Z digest=sha256:6506317702e755198d1f739e683fe738a3504641655df95efb9bf16d19e1f172

Observation e021a51d-e594-4c22-9eb9-a07b0d2b105e · outbound

This paper cites Contrastive Learning with Hard Negative Samples.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Contrastive Learning with Hard Negative Samples

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:43.344779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:43.344779Z digest=sha256:ffc962b126520788421a18996ed65c03d274b7685d278549b9be8207c8677a25

Observation 465889c8-cf34-48ac-9603-90480d251467 · outbound

This paper cites Pic2word: Mapping pictures to words for zero-shot composed image retrieval.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Pic2word: Mapping pictures to words for zero-shot composed image retrieval

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:54.171773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:43.400909Z digest=sha256:408696c903ccfe8b89cf9ae42775bac5a9d039ff239b80a357c29ee08309e7cb

Observation 9c0da33c-6771-4f97-be58-70b508e1a057 · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Laion- 5b: An open large-scale dataset for training next generation image-text models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:53.924513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:43.456451Z digest=sha256:e6d7db9f15a496adf35904be028d091ee9627a3374dd36361c608c83f6bf8d0d

Observation 410094f7-2180-4340-a721-0811d9523df3 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval A-okvqa: A benchmark for visual question answering using world knowledge

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:53.690728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:43.506063Z digest=sha256:3a6f552ae66a0d0ea3a8e88535ee8f9530773739f0990c7f78e0228be7036399

Observation 1393e73d-2486-4a8e-9db2-783e9b418355 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:43.554578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:43.554578Z digest=sha256:6b5aca1f0baec8a68e2b528507bd96478aaaf18a95d68416fef2d3cfbb1ca2c1

Observation 4fc5fba1-6b55-444f-83ac-485e825a272f · outbound

This paper cites Composed video retrieval via enriched context and discriminative embeddings.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Composed video retrieval via enriched context and discriminative embeddings

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:53.527330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:43.611833Z digest=sha256:44b1dcc06a46ffe1adddd15f28b1ee5c3f07fd9f5493d78bc7c3aaae36f42248

Observation 9ddbad45-f0bd-4905-afec-10a0b6c37f36 · outbound

This paper cites Fever: a large-scale dataset for fact extraction and verification.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Fever: a large-scale dataset for fact extraction and verification

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:53.281847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:43.665331Z digest=sha256:f66eba52b72626ec1b575a7563c70fc3864f1022b06c845d33398f851051c5cf

Observation 7e49834f-8235-49d4-ab1a-589f9b41a9d2 · outbound

This paper cites Covr-2: Automatic data construction for composed video retrieval.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Covr-2: Automatic data construction for composed video retrieval.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:53.036132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:43.745704Z digest=sha256:97a4f63dc62f754a4188fed664bddf55ef9f47bc31271cb8dd85df62bf4c93a6

Observation 43ac363e-3f70-45b1-b50d-cb55d8b92e58 · outbound

This paper cites Covr: Learning composed video retrieval from web video captions.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Covr: Learning composed video retrieval from web video captions

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:52.799218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:43.885501Z digest=sha256:69d98aecba93b84cebe30dd3f1a93d5019ea8b897146f8b011d98dd345233edf

Observation 4165373e-c41e-4f73-bb44-aa18251c6109 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:44.083171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:44.083171Z digest=sha256:54d1d61c6ec2701c1ba56eab57d934219fc1c34df0df630c8369881e27db2526

Observation fa90b388-4267-4933-acb7-be02d1deee11 · outbound

This paper cites A Comprehensive Survey on Cross-modal Retrieval.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval A Comprehensive Survey on Cross-modal Retrieval

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:44.181075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:44.181075Z digest=sha256:b68cfb1497c280d3ef3a2e79b958267c0c9f0d91629d66c4c54b2ddf081e9a4c

Observation 112d2fbc-dfce-415c-9cee-36e5875b4c65 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:44.311504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:44.311504Z digest=sha256:2332f1cbd1c210ce4d7411e50e169a7b181f38bf24e1ac1ffd23aad9ae17c1b9

Observation 00e1ee86-7f82-4c3e-9637-2ac103fc5528 · outbound

This paper cites Fvqa: Fact-based visual question answering.IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(10):2413–2427, 2017.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Fvqa: Fact-based visual question answering.IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(10):2413–2427, 2017

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:52.620569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:44.452960Z digest=sha256:a123239f693293b3536bb77132c3eac25b54a5de0f8f29d1d5f0741de8cf0e14

Observation f7e16bdf-8b49-4887-aea7-5fe6a581e0f8 · outbound

This paper cites Cross- modal retrieval: a systematic review of methods and future directions.Proceedings of the IEEE, 2025.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Cross- modal retrieval: a systematic review of methods and future directions.Proceedings of the IEEE, 2025

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:52.402528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:44.602734Z digest=sha256:1ddb418b673fb66b08144a45d1543269ea91daa7f68f93c41a823d3478bc481d

Observation 22241213-305e-4750-a912-386dc00fafc9 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:52.210759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:44.751774Z digest=sha256:15f1aee4f4c768ab4e935dc3c2d8c6e377d89276b0c5128838289e8015847c2e

Observation 1ccdfecc-be1c-4081-870d-5ebf432c8179 · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Internvideo2: Scaling foundation models for multimodal video understanding

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:51.966932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:44.896788Z digest=sha256:1352181a07da05f8ec0bca3feb9bc13d3e654e6885b299e1ab39628c7f3c6b9a

Observation c92fe79b-13d1-4bc5-a59a-c97df7637a2d · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.013607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.013607Z digest=sha256:fa8586d2e4eb954c1df08384d3e20218b4987fba4b241f11b42e8f9bf4fffdbb

Observation f54c8e76-bdc0-480b-9fc8-a497a0ec08de · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.104101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.104101Z digest=sha256:7ee1721a4eb5fbc6612a635de1afde6833961fa01c295d05267e90f144adc732

Observation e20025dc-ed00-4245-aed1-d7b7f54c7576 · outbound

This paper cites OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.187968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.187968Z digest=sha256:4aa955dcaff5d2cfaa237b50be1b1db539feb246f670d85cf764286deb0c272c

Observation 4dc596e4-9b13-4caa-9b6a-29a932e200f5 · outbound

This paper cites N24news: A new dataset for multimodal news classification.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval N24news: A new dataset for multimodal news classification

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:51.794730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:45.273016Z digest=sha256:de310a3a95d19f39b38755783e4a4a7683d49504a42c0b65568cdc45bcd83f6c

Observation 2c6dc2a6-8756-4a3a-bc80-de2df8aaa6c1 · outbound

This paper cites Uniir: Training and benchmarking universal multimodal information retrievers.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Uniir: Training and benchmarking universal multimodal information retrievers

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:51.598369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:45.369387Z digest=sha256:9cb6925a58bc7b75e57ce0036757b4b1c043d8c74065b3fd4f6ec41c40f175e7

Observation 0ccbf610-017b-47cf-a0cf-93ed9d61fcfd · outbound

This paper cites Sun database: Large-scale scene recognition from abbey to zoo.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Sun database: Large-scale scene recognition from abbey to zoo

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:51.405208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:45.437690Z digest=sha256:f74f924c3a0a65036ac264dd7598baf03b71cbdc31f22c733d88167b750b8cfb

Observation b10f7c85-ecde-490e-acd6-857b03f829f8 · outbound

This paper cites Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.558531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.558531Z digest=sha256:8de404842b2896a7a9e6dc6babcf4471477acef8b4521ada0558d1e4f56b503c

Observation fafc94c8-f7eb-405c-bbee-bdbf296ecd29 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Msr-vtt: A large video description dataset for bridging video and language

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:51.233064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:45.631311Z digest=sha256:fd1fbcbb89899c86eef6337325a5c00dc3af4a57038fb1ea34d308e4bf5e2f62

Observation 731ad417-6323-44dd-8cd0-da784f786d16 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.703524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.703524Z digest=sha256:9f7343c95f485d18bebb33a1516b97a5abf62ae90cf77b7aaac05eeee6892718

Observation ff5b82ca-a31b-43ae-aebe-5614bf80505d · outbound

This paper cites CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.772749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.772749Z digest=sha256:5968686655db51cfb3577661395b4dc804b12b170ee8c12858f46c51d1b57ca8

Observation 5f71f51c-5bf3-4289-8e15-5670d5225064 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.839239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.839239Z digest=sha256:fae4ab1550122a0dbaa99eb5b125bee753a86ae21d58068b9aec91bc107527f3

Observation 0da0f99f-033f-4503-b951-6c6d25f52ded · outbound

This paper cites End-to-end multimodal fact-checking and explanation generation: A challenging dataset and models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval End-to-end multimodal fact-checking and explanation generation: A challenging dataset and models

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:51.025254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:14:45.930404Z digest=sha256:bd58d17c86d7e8aad143d2d0d1c3ad81b5677a25692acd319a13c9ee77d3603e

Observation 0b5c142e-69e9-4709-91ce-7bed037176bb · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.015018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.015018Z digest=sha256:84f09bd0681233cc31111cbe2f51dddeb8d072049d1eec536275cca19b6ce54a

Observation 25d936f7-e057-4d04-b373-1597cff5eec4 · outbound

This paper cites CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.090733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.090733Z digest=sha256:9f2ffe3dff4cca7c9375f7f3e3fa709361d12dfc61139dbd0e6406892a8b5cfd

Pith citing papers

Observation ee1ebd9e-a9c9-45ae-ae6a-38cd48861805 · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:00.563577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:00.563577Z digest=sha256:2eeb2e82fd4051e1067031c3aa1f9d2aab66d6f591353df9836e52046cd26693

Observation edbdef5e-892a-46ba-bff0-03e3b5d5177a · inbound

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction cites this paper.

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:11:27.464856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T14:09:22.942238Z digest=sha256:a58f368489cf7c263270780bb2ac376f20e204eedb56cb97de1a8e10ca512b50

Observation 913342ff-b8a1-4a73-8d70-b9dd8edb213b · inbound

FreeRet: MLLMs as Training-Free Retrievers cites this paper.

FreeRet: MLLMs as Training-Free Retrievers Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:01:23.532492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T13:00:31.952588Z digest=sha256:8e1de767ae504739a83def96ca1491530261bf873625294b5de42e1debcc0989

Observation 53c0bd72-bc9e-42d9-87d1-38ee6fc94016 · inbound

FreeRet: MLLMs as Training-Free Retrievers cites this paper.

FreeRet: MLLMs as Training-Free Retrievers Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:45.017918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:51:45.017918Z digest=sha256:fb86a820400986ef5595cb5acfd8cf6e41804cc5fa457806230f384c69424ba2

Observation aa897247-7e0b-4148-9f23-44169316f1bc · inbound

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings cites this paper.

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:01:19.736570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:46:30.827346Z digest=sha256:aacc9ee1164880fa90ec65e1fa368a9748fdae12c19f0a887e31d4dccbac3984

Observation 8f2ce394-f8b1-45e9-893c-e0de3bc7989f · inbound

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings cites this paper.

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T18:31:37.144968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:31:37.144968Z digest=sha256:4bc85c69d84c5e202027fd41afd9fad13e6a275720221226458302e057ab5897

Observation 1c769451-754c-4c00-a1c5-adf5685cc380 · inbound

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings cites this paper.

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T15:40:29.161499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:40:29.161499Z digest=sha256:f537e81e4ab08ff3f53934d37d7d96f43319d5ccc7471817acca56dfde272fa1

Observation 8a9ccc3a-2e9a-4e11-9033-1cd3626b5ee9 · inbound

Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval cites this paper.

Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:13.551343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T16:56:52.714346Z digest=sha256:074ac5bc301755c3d4f1280db31c80e1c62f78e69600dd53343b47e0deb53b4a

Observation 8a5273d6-52b2-4e9e-9818-5a94e51accd8 · inbound

SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval cites this paper.

SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 5

Resolution
malformed identifier
no resolver link, observed 2026-07-14T18:56:57.352952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:56:57.352952Z digest=sha256:1de0a28d4f10aadc46b5c5adebc7bb8fec1a40d7ec5bd5aa386b34e8cd7162f1

Observation 19bd5014-640a-4ada-a4d1-8c375f21155c · inbound

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval cites this paper.

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T05:49:36.604036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T15:34:55.062016Z digest=sha256:21976e5128edbc36a0cf7d600ca6f7899081fc46814f453e4413acfc186940ec

Observation 68e4c1a0-4d0d-436a-b85c-8bffd421c7ac · inbound

Illuminating Visual Identity in Universal Multimodal Embeddings cites this paper.

Illuminating Visual Identity in Universal Multimodal Embeddings Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.106644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.106644Z digest=sha256:6db115af74fbe718b7f4c2aac7dfbb6cbd03f61d18b6b3c10bca4a59da55761c

Observation 8ba9a618-5492-42cf-be31-7c0106c94e0c · inbound

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval cites this paper.

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T18:30:14.559722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:30:14.559722Z digest=sha256:7801971efcd99c8f25ec7785796348a85d09450d10b29446e7c4f6253c52b904