Pith. sign in

Paper Citation Record · LEDGER

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs

As of 21 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 0 inbound Pith citation observations for arXiv:2602.05275.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.05275 v2

Coverage vector

measured 100 of 109 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:20:59.229137Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 109 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 166a2c47-85ae-4c04-906c-9eee29fe86a9 · outbound

This paper cites Qwen3-VL Technical Report.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:50.269672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:50.269672Z digest=sha256:076960548c1c8eb19f077be4bb6d5333262f2bec3e591c9cf53f6ac420e7ae46

Observation be59d751-9b27-45f1-93fc-590044a5f705 · outbound

This paper cites Coig-cqia: Quality is all you need for chinese instruction fine-tuning, 2024.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Coig-cqia: Quality is all you need for chinese instruction fine-tuning, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:50.328330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:50.328330Z digest=sha256:8ae4db3156d2d33e5cb83cd5de20b379dbc97bdde67e3c5714ae2a9bd55c4cad

Observation 05d1e918-7453-4f08-aeda-dba483f41331 · outbound

This paper cites Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models.Advances in neural information processing systems, 32, 2019.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models.Advances in neural information processing systems, 32, 2019

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:50.405091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:50.405091Z digest=sha256:f1c11fcf4c423ee39a673d3a49e29d7b87ff8eedb3f71ac1c09499bfb9a0c36a

Observation 8d3a293f-257b-4725-9561-cd6f6ba388f5 · outbound

This paper cites Baai-mtp dataset.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Baai-mtp dataset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:50.489896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:50.489896Z digest=sha256:d76a06bdce5db90a91899a8b4c71ac43b1044ac302263de6f46a2804f0df35ed

Observation dccbc202-eb71-431c-a90a-e929ada7b54e · outbound

This paper cites FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:50.578691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:50.578691Z digest=sha256:047a8e971701377986c2fc5284fb19232150b0636d2a593206ae0c66dc525986

Observation b69c64be-4bfc-4880-8c75-5a0d9ff684d0 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Honeybee: Locality-enhanced projector for multimodal llm

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:50.649295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:50.649295Z digest=sha256:bdb204647a49f5fb6cafed1c3b120288d793111ed2816564213dd5b75f20cae0

Observation 7868e2cf-7529-46d5-b645-eb23bab88664 · outbound

This paper cites Webqa: Multihop and multimodal qa.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Webqa: Multihop and multimodal qa

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:50.726553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:50.726553Z digest=sha256:f4e6b31e604338a725c2123854792241bede2c35b561bb2df0f2ba1cfd7c2db5

Observation 9c243006-424a-4db2-b9c6-014a5ba99d4d · outbound

This paper cites mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:50.826332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:50.826332Z digest=sha256:4cc74fdea329da5ac0feb76b5f90ebc77478d0e03adf1b030f80586bf4999eda

Observation 2335a39d-00c1-49f2-ada7-3d9eccb47a02 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Sharegpt4v: Improving large multi-modal models with better captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:50.986221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:50.986221Z digest=sha256:6b2061ce6cf6a9ba5aa958ab81f4648772d9955c1552cd302b3e2e33d8776d6b

Observation acab2e52-707d-4483-a256-c636ab96fd9b · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Reproducible scaling laws for contrastive language-image learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:51.146796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:51.146796Z digest=sha256:0e847915ae901465520461b13d0bab9b08bfd61d280fdf3c95a7d4c30e512e24

Observation 7ce5ec27-ebf6-4054-8c4b-ab3e8b33442d · outbound

This paper cites Think then embed: Generative context improves multimodal embedding.arXiv preprint arXiv:2510.05014, 2025.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Think then embed: Generative context improves multimodal embedding.arXiv preprint arXiv:2510.05014, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:51.273014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:51.273014Z digest=sha256:4c69f7199d15fb4be8b691f371a1f75bc732d22dcbc8dc4f0c3d9298a2dc9ec0

Observation 24da3b8c-1f04-49fd-abd6-26b21d0ed7f5 · outbound

This paper cites Visual dialog.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Visual dialog

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:51.393186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:51.393186Z digest=sha256:d6ebfb8807682d53f509828b7bdff19b2fd1528fba5a0512abb9474e875c74c5

Observation 3afd858a-3cbb-45c8-9c88-1b34042db5d1 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Imagenet: A large-scale hierarchical image database

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:51.514888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:51.514888Z digest=sha256:ffdc7f1d021a6130a9d8b1e82d9278a057428d5b00e929fc70c56e7d4abeafaf

Observation 48688fd7-db64-449c-a223-2b57973041c5 · outbound

This paper cites Bert: Pre-training of deep bidi- rectional transformers for language understanding.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Bert: Pre-training of deep bidi- rectional transformers for language understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:51.658647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:51.658647Z digest=sha256:a386e11377a613677813d4dc4b7c8e1e5c51d69adc652eb8eca22669e074eba6

Observation 179adbdc-b9d6-47a0-a487-8597ebc8b8b6 · outbound

This paper cites Pact: Pruning and clustering-based token reduction for faster visual language models.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Pact: Pruning and clustering-based token reduction for faster visual language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:51.796555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:51.796555Z digest=sha256:63dd8b8b12b705d716ed2d289b3ffb45122203e623d564332266b28b78e37f56

Observation 31325f37-7599-413e-91ee-3397c878c3b1 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs An image is worth 16x16 words: Transformers for image recognition at scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:51.902230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:51.902230Z digest=sha256:6749b761f7be7e637dacab79e6979a75c5eb959aef8d42c61d780b2ee4844b41

Observation 43981ea6-2a2e-434c-a555-09f7d8b36f02 · outbound

This paper cites The pascal visual object classes challenge: A retrospective.International journal of computer vision, 111(1):98–136, 2015.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs The pascal visual object classes challenge: A retrospective.International journal of computer vision, 111(1):98–136, 2015

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:52.014405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:52.014405Z digest=sha256:021c814089899ddba56d54baf34155953de9aca2b0c3c1172bdb4fd3ce31af7a

Observation 46bf6137-d634-43fd-a7ac-5214bd75eebb · outbound

This paper cites ColPali: Efficient Document Retrieval with Vision Language Models.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs ColPali: Efficient Document Retrieval with Vision Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:52.116209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:52.116209Z digest=sha256:cc68d65a6273a0db5724e2ee8b60a398454930df97d2503576d39a885b88c9ca

Observation 183a65f1-5665-498b-9cbf-7a05848fb385 · outbound

This paper cites DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:52.280678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:52.280678Z digest=sha256:e4146da2389b3fab99910daee91d303ef6f714140e2ff0edd9568ca9086091a2

Observation 36da3974-7e5d-4c11-8b8a-6abe816d0e2d · outbound

This paper cites Vl-clip: Enhancing multimodal recommendations via visual grounding and llm-augmented clip embeddings.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Vl-clip: Enhancing multimodal recommendations via visual grounding and llm-augmented clip embeddings

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:52.411606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:52.411606Z digest=sha256:b521e40a332dddd34c4574efc55b70cce24bb19ea34341424154db166e47190c

Observation ae888698-4138-4df4-bfa0-ca4c4b4bc9d8 · outbound

This paper cites Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:52.540683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:52.540683Z digest=sha256:74d7b2d9a2b01c686cc3e585e3818237124abdc1e6e926be7c319104bfc52471

Observation 6d04b80f-cadd-4690-bdf0-76de269deeb4 · outbound

This paper cites Breaking the modality barrier: Universal embedding learning with multimodal llms.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Breaking the modality barrier: Universal embedding learning with multimodal llms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:52.677077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:52.677077Z digest=sha256:c7ab64a34f22e7ff1016443ec3ca7e960412d22a3a54f71476d279225f5bed13

Observation 09dd4733-71ab-4e64-9960-30af4ee00814 · outbound

This paper cites Unime-v2: Mllm-as-a-judge for universal multimodal embedding learning.AAAI, 2026.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Unime-v2: Mllm-as-a-judge for universal multimodal embedding learning.AAAI, 2026

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:52.837118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:52.837118Z digest=sha256:ec754b842a9154baa8cdc76b961fe4ad12d2680f564cd491b65e1dd9f72263c3

Observation fa8ff887-223a-4a26-965b-f40dcfcec6f0 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Vizwiz grand challenge: Answering visual questions from blind people

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:52.965109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:52.965109Z digest=sha256:a9447a3454fb5e7c7aaec2cf39cbd566161c88c84dad5be6b39b5a61026141ab

Observation 82da0299-3d75-4c2d-ab49-b347be8f59b6 · outbound

This paper cites Efficient Multimodal Learning from Data-centric Perspective.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.053118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.053118Z digest=sha256:62f453699d602c6b0939da178fec107541373a8057d5843f2e7ecf9773c9de7f

Observation 96a5728c-d82b-42ed-b2c6-63046e03b7c3 · outbound

This paper cites The many faces of robustness: A critical analysis of out-of- distribution generalization.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs The many faces of robustness: A critical analysis of out-of- distribution generalization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.141125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.141125Z digest=sha256:e898f7a54565fd2a0a5e0f97a86ce294c34299687d37d46714668a9f21bedd12

Observation d99ed3e6-fe96-4b1b-8de8-809fbf6fa007 · outbound

This paper cites Natural adversarial examples.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Natural adversarial examples

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.211095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.211095Z digest=sha256:afdb0a3781c2c834e9bf52bf9ba84105997b77b4069f2a09e186617adb780687

Observation 5d802671-6139-46ae-a30b-c63e2d2b8bd8 · outbound

This paper cites Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36:31096–31116, 2023.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36:31096–31116, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.278887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.278887Z digest=sha256:e4db71b92a33c16f594ecbbd0965015fbfd34237b40c8b3183722e8d07734f69

Observation 91081082-ff54-49d8-b4d1-911425dcc288 · outbound

This paper cites Open-domain visual entity recognition: Towards recognizing millions of wikipedia entities.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Open-domain visual entity recognition: Towards recognizing millions of wikipedia entities

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.357148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.357148Z digest=sha256:b33d494440e4046c41a8938391e85bcf4da7f1c44917fb511d6385db6427cc6b

Observation cfb612ed-935a-4f4b-80d7-b1570230c04f · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.430200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.430200Z digest=sha256:09892c239d331cba00885984e85e0edbf41330a0f4ec41277d673a83f170de35

Observation 8bc95e34-6faa-4aaf-8c14-fbb97c7c3366 · outbound

This paper cites VideoRAG: Retrieval-Augmented Generation over Video Corpus.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.503874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.503874Z digest=sha256:30202c308c792d060a804f78f49a603de94252b363bc9ab56e0516cd80e5d7fc

Observation acb3f71a-b666-48ae-aaf1-634564865cbe · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Scaling up visual and vision-language representation learning with noisy text supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.580062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.580062Z digest=sha256:c980b547a65bc323f8381711f0f185f8eaa85b124409e443cfb7315cdafb6737

Observation 977c4326-5f16-48ca-b8b6-6358a3d87dcf · outbound

This paper cites Rzenembed: Towards comprehensive multimodal retrieval.CoRR, abs/2510.27350, 2025.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Rzenembed: Towards comprehensive multimodal retrieval.CoRR, abs/2510.27350, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.640896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.640896Z digest=sha256:8f33cb698886d7008af0662d333f92608b338016008a4e219ed9e4d9d843f7c0

Observation d9095c6f-3ea3-4a98-b77b-b176b4c5dc82 · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.698736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.698736Z digest=sha256:545e100efbaa7756be7370a5e62cb0166f52d6d21274dc50a9914e09eb718ea5

Observation 12f5a73a-ca73-436e-acd4-ea478d6def41 · outbound

This paper cites Vlm2vec: Training vision-language models for massive multimodal embedding tasks.ICLR, 2025.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Vlm2vec: Training vision-language models for massive multimodal embedding tasks.ICLR, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.795182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.795182Z digest=sha256:c7133450020ef9c525d475cd8eb415093ef8495aa5190888097e28744680c1c4

Observation 25b78c6e-7297-431e-842c-9149d7217160 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Referitgame: Referring to objects in photographs of natural scenes

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.858465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.858465Z digest=sha256:1ab3d9d0732b1e0d553b66a05de09ae3cfdca7ab1ea51c802e2030bd6a0b4c94

Observation dd25e9aa-bbbf-412f-b772-e53181d135bc · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.Advances in neural information processing systems, 33:2611–2624, 2020.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs The hateful memes challenge: Detecting hate speech in multimodal memes.Advances in neural information processing systems, 33:2611–2624, 2020

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.910519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.910519Z digest=sha256:337e9f43d013fe1884e634b393d7044ad9efdc0bfb79a972299a927c28feb5f7

Observation f8b59236-ee56-45db-b021-deaa1d229b16 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.976937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.976937Z digest=sha256:34751a8c94078c9be7e11e5936ad0477d25228d505adefec0fceaa78b3444557

Observation 9ba3d8cc-e5eb-4ab6-b1c9-027339cbd6c6 · outbound

This paper cites an unresolved cited work.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.032303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.032303Z digest=sha256:44ce7157504eebd4173724b75d0bfc08b2387b0c31c301f2c1b6954e6b8f18e5

Observation 213c22ef-aa82-46dd-97b9-45934dc05b11 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Gonzalez, Hao Zhang, and Ion Stoica

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.086589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.086589Z digest=sha256:b008f6fbbfb95ccb7c39eb22c414de233cf7c6756a9feebc038f2e7e496d7119

Observation 7c4ddbba-ec11-4314-9fdd-3a0ff0b81ff3 · outbound

This paper cites Llave: Large language and vision embedding models with hardness-weighted contrastive learning.CoRR, abs/2503.04812, 2025.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Llave: Large language and vision embedding models with hardness-weighted contrastive learning.CoRR, abs/2503.04812, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.147896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.147896Z digest=sha256:15070bc12a06b518d6662b456b5de0972bf5e779687c9edc5155d30cfcd5a8f6

Observation 929d70e2-7b2a-48cc-a0c1-b4a09760ae73 · outbound

This paper cites Ume-r1: Exploring reasoning-driven generative multimodal embeddings.arXiv preprint arXiv:2511.00405, 2025.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Ume-r1: Exploring reasoning-driven generative multimodal embeddings.arXiv preprint arXiv:2511.00405, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.207577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.207577Z digest=sha256:06403254b0f39b9ba25aca93a93caf1156879eaabc21912543fb325af86b363e

Observation d1a1b7cc-4327-4d06-8e40-7aaa6efa2abc · outbound

This paper cites Building and better understanding vision-language models: insights and future directions.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Building and better understanding vision-language models: insights and future directions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.248497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.248497Z digest=sha256:2ad80399cb850abf44b600613df7b4624810475446116fa10774a6e434503e92

Observation 8737e8ec-8fc9-461a-97ec-70b28ff7d280 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.294979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.294979Z digest=sha256:f6b671da8232e25082d92b78169f12d16554a1942e97320f157d0686aa7a7353

Observation 28e058cf-59cc-429f-a3e6-6b3e1948a423 · outbound

This paper cites Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.360576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.360576Z digest=sha256:7152fcdd463f8a1dfa29e67964b15066df041b1682c487d8410e92164faec2c7

Observation 2b1f1732-8e9b-4c7c-97ae-e17119a92bef · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.428984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.428984Z digest=sha256:e5bd522d05d5ecdf0d4cc5bc86bfe1124f6ca449ea1ca0fed8a49992c4234253

Observation 1f1bd9c3-28b3-4f0a-9308-6f87286f58e3 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.527590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.527590Z digest=sha256:a345aa481982e75448b85113cb0c2b3724de0afaa4936c2f1278577388dfa328

Observation bcfcfc53-0e48-4a26-963a-ef15269b830b · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Silkie: Preference Distillation for Large Visual Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.615173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.615173Z digest=sha256:35262fefc394489b52cd4c72ae3704c956e680f98d44edea369b36f2ef3da2c3

Observation fc433143-8e98-4778-8378-95c28f42aceb · outbound

This paper cites Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.697053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.697053Z digest=sha256:bef5c7f41f2cf278f38c7afcb8e3ad5727acf626fecd3e33ce1e781667ca441b

Observation f805cdb2-ece6-4b0e-ad8e-6766ace1d552 · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Monkey: Image resolution and text label are important things for large multi-modal models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.749572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.749572Z digest=sha256:d411c9301da5aa71c8b0704aea0637105f47ec2ec061c0c7b0182ac99d782b2b

Observation dd0e4e57-7762-4cbd-815f-a76f08928956 · outbound

This paper cites MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.793364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.793364Z digest=sha256:caaca9f5fb2618850d52c6ec9352b5c1289947f69c79446c67327452d512ab23

Observation 8c2cb28b-4ad8-41a6-9bed-55a112eca2a0 · outbound

This paper cites Microsoft coco: Common objects in context.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Microsoft coco: Common objects in context

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.860381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.860381Z digest=sha256:cf2fbffa2c44d1d6401d5227e6820b4e0e1abad00c80232d9cdfa794d79d4b4d

Observation 9143fca5-9bfc-428a-a83b-e062fc6b8bf0 · outbound

This paper cites Gres: Generalized referring expression segmentation.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Gres: Generalized referring expression segmentation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.915681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.915681Z digest=sha256:1f72bd98f232effddcde3c8a6d935a58ad0c34032060e85d4010a0457ba914f9

Observation 190b143f-5bb8-4e52-b0fd-96324161291e · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.954439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.954439Z digest=sha256:ac1b7d73b4a56e7be5192566a9e8ac9757ec6c7efefc4b28d34f451fce7044af

Observation 2de17b5a-c7b5-472b-9aba-918622fe9e6f · outbound

This paper cites Visual news: Benchmark and challenges in news image captioning.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Visual news: Benchmark and challenges in news image captioning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:55.043738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:55.043738Z digest=sha256:236387355baa4dce885709fe7f1af907e90e7a8566e6f6437bba4ced3af39b5c

Observation 12f91ffe-ddfb-400a-bc06-bee301d4359a · outbound

This paper cites Improved baselines with visual instruction tuning.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Improved baselines with visual instruction tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:55.111831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:55.111831Z digest=sha256:23311a4a0f777188d96c5bfdfc977bc811c7ae3a10f73d5f6b14e41da822418b

Observation 7f072e3e-75bf-43f4-a3e2-bcaf95f242df · outbound

This paper cites Exploiting transformation invariance and equivariance for self-supervised sound localisation.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Exploiting transformation invariance and equivariance for self-supervised sound localisation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:55.197386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:55.197386Z digest=sha256:0af116ef7721ef5f1ff2fae436ff2afab17f90304d430e230b11bbb8df5ea709

Observation c2c6451b-e443-4568-9f64-85f32a287c39 · outbound

This paper cites Edis: Entity-driven image search over multimodal web content.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Edis: Entity-driven image search over multimodal web content

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:55.280158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:55.280158Z digest=sha256:dac8f2f43ef30fe1691bed183a11b4ae184c206477db52e6c73a6ab8a9996480

Observation 00fcc7cc-c367-458a-93c4-218c97b71f19 · outbound

This paper cites Lamra: Large multimodal model as your advanced retrieval assistant.CVPR, 2024.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Lamra: Large multimodal model as your advanced retrieval assistant.CVPR, 2024

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:55.374183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:55.374183Z digest=sha256:7f6c65d9ff11b046adec966ff0365db8bb3242860fc2142173c115fa945aecc9

Observation 95a0d4ea-cc76-44ac-9e27-733a7b3bd8cb · outbound

This paper cites Image retrieval on real-life images with pre-trained vision-and-language models.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Image retrieval on real-life images with pre-trained vision-and-language models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:55.501729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:55.501729Z digest=sha256:226c06360428879a4e75fd1599a0718b34383bc2a9fefcc5f7c823a8fe4d9135

Observation a45aa92c-f61c-4a21-931f-c62c279a5b61 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:55.620998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:55.620998Z digest=sha256:d6818ada35e11392da9a4b24a0404254d9ff31d2c08c89864ca445a5c3c835a5

Observation 58ad22d6-b653-4236-9404-83176f58b5ae · outbound

This paper cites Unifying Multimodal Retrieval via Document Screenshot Embedding.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:55.730245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:55.730245Z digest=sha256:8d45a9ac35d60f9caff508c1a1e70775ab7360472db2eb39c615f7c6cacde60c

Observation b18985b3-4092-4a88-8509-22269840a629 · outbound

This paper cites Visa: Retrieval augmented generation with visual source attribution.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Visa: Retrieval augmented generation with visual source attribution

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:55.841007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:55.841007Z digest=sha256:08306569fff4dfa6018218363891e03b4fed7c2cbdb1f4a9e792369addca3a1e

Observation 86b9a751-c795-4086-9ac3-21f8e45aeff7 · outbound

This paper cites Mmlongbench-doc: Benchmarking long-context document understanding with visualizations.Advances in Neural Information Processing Systems, 37:95963–96010, 2024.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Mmlongbench-doc: Benchmarking long-context document understanding with visualizations.Advances in Neural Information Processing Systems, 37:95963–96010, 2024

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:55.954787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:55.954787Z digest=sha256:0fd6401da066cf3995e33b72e9c9f5b4946ceda7c82787686611be010f292680

Observation c34fcdb9-e9d0-496e-b7d3-556391baf8b7 · outbound

This paper cites Vidore benchmark v2: Raising the bar for visual retrieval.arXiv preprint arXiv:2505.17166, 2025.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Vidore benchmark v2: Raising the bar for visual retrieval.arXiv preprint arXiv:2505.17166, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:56.034599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:56.034599Z digest=sha256:be9b0eac4e42157f007d43465531c11b4298b2f56e0943d7a62df521dbac5a15

Observation 7064aa75-1c13-4bb3-b89f-41d8c54ce929 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Generation and comprehension of unambiguous object descriptions

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:56.106545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:56.106545Z digest=sha256:8bdf0cb8f24250a21f5eb9f4e22da30c26011b102ff08dc109cd92f49d9b27e5

Observation b8adf7aa-019e-4dbc-b2ac-7fc6aa70587c · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:56.179526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:56.179526Z digest=sha256:20d152d9e63c4cea5bd31dfadbfd16915241b98d9702c1eff4e09f5f8a5fb093

Observation c3f7e06e-2efe-48a1-8021-d91f584d35a6 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:56.278664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:56.278664Z digest=sha256:d8fc0154d52e1cfcc85790aa91c1de8faaa6f794509a689d715512d2fc1c2117

Observation 74284f40-3807-4964-9064-ede3c1dffadd · outbound

This paper cites In- fographicvqa.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs In- fographicvqa

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:56.348173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:56.348173Z digest=sha256:4ebb1a2dd1d6e0f5a4e330c02983e48be13bed97a3ea605bf9d1e5417964f38f

Observation 0468aa6e-2419-4b77-92df-d4c586a2df12 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Docvqa: A dataset for vqa on document images

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:56.454771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:56.454771Z digest=sha256:048f54265cb427321b73a7d6488254eec34db4ccc7fee364a3ae701e56d3162f

Observation 5b9ed98a-34ad-43c7-a485-f8ac7fb784f7 · outbound

This paper cites VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:56.545981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:56.545981Z digest=sha256:381855dcdcb0719fb924abf45b999b3168df7aa270b29555ad659842a6b35462

Observation 07739cff-4fdd-40ed-a91e-678f06fad60a · outbound

This paper cites Ops-mm-embedding-v1.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Ops-mm-embedding-v1

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:56.651977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:56.651977Z digest=sha256:493d69181e94cab48cd82f6ab6a0f1df77027c909a79e4fa0152bd959841b187

Observation b96811c7-27b7-4610-86cc-2b0c40849d8a · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:56.773017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:56.773017Z digest=sha256:590ed83635bc4891926ee2ad9a7cbe51396555777ce47d131bd14225d5019c62

Observation c18fb213-721c-4e00-80c4-94d9f798838a · outbound

This paper cites Learning transferable visual models from natural language supervision.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Learning transferable visual models from natural language supervision

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:56.860259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:56.860259Z digest=sha256:9ff0bef9558dd9963cb4e63bc03a1a1e31307455dba2892d9ee9a78aaebc6424

Observation 7d885533-b961-4aff-ac6c-87c1f239c3c9 · outbound

This paper cites A- okvqa: A benchmark for visual question answering using world knowledge.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs A- okvqa: A benchmark for visual question answering using world knowledge

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:56.955766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:56.955766Z digest=sha256:ae89939971dd04c854132e985b67193360aead88a9ee53c1e4b41c04ba470dbf

Observation 004919ca-0b3c-4749-9a9e-1e8998d48a18 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Objects365: A large-scale, high-quality dataset for object detection

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:57.082636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:57.082636Z digest=sha256:96246639149c8d25eb0d5d3fe2c5a7fbfd82d63458f5335bfb9ba3eaa3085240

Observation aafa2c93-40bf-4735-a64c-b8f7d18dbf5d · outbound

This paper cites Sharegpt-chinese-english-90k: A bilingual chinese-english human-machine dialogue dataset.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Sharegpt-chinese-english-90k: A bilingual chinese-english human-machine dialogue dataset

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:57.170835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:57.170835Z digest=sha256:8ba0be6e922d71c9b7e9571030a1a231d57f8b4a46221749d60d0d84dc90cf53

Observation 1e76efb8-36ed-4ea1-93e3-bfc2a5ee58ce · outbound

This paper cites Towards vqa models that can read.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Towards vqa models that can read

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:57.249296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:57.249296Z digest=sha256:bf76d9226de0be14548621a85d79eeb5202e353014f03b64a74208f1beae3d2d

Observation 24b1a7b5-fabe-48c9-99c0-5ccc3d4e1fe9 · outbound

This paper cites EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:57.341823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:57.341823Z digest=sha256:0b11caa70ca632f24dac2c2829c73af7c00df334950df9e9cbcd63db9be29984

Observation faabcb24-c86a-452e-a640-9a27f8a5cd1d · outbound

This paper cites Breaking the batch barrier (b3) of contrastive learning via smart batch mining.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Breaking the batch barrier (b3) of contrastive learning via smart batch mining

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:57.407098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:57.407098Z digest=sha256:834cc68ebb9b2228eda89120d2a695098c7e15428b2fb5d8fe58698c3b2f1aa1

Observation 8fe26576-c358-45b5-a0c1-e3d8593f75b0 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Representation Learning with Contrastive Predictive Coding

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:57.478991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:57.478991Z digest=sha256:03ada7bdd0af8cd2e0b77abc94fbe1718821ecc22450dd155eb76d0f7f1b9262

Observation 596c61a7-975b-44b2-82cc-8b4996b6aef3 · outbound

This paper cites V3det: Vast vocabulary visual detection dataset.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs V3det: Vast vocabulary visual detection dataset

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:57.579567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:57.579567Z digest=sha256:f0a0abd9ab740e8de741e4613d8a9a0d487332731ac814d11b6519793b84b9d7

Observation 260f681b-8edd-4671-8ead-82b65f1cc844 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:57.669229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:57.669229Z digest=sha256:37ea5f16cba72b2db5bcf9efcb7bc1f45c45496ce33014b3ec8d8a0694987dd7

Observation abafe370-d164-4aaa-a5ea-aae9a7ffed09 · outbound

This paper cites ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:57.766049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:57.766049Z digest=sha256:1dac44a14e5a44ae30757e9a6c6a71142f1e83d5ba5cd41e3e5ebffef59517c0

Observation 5b61da42-886d-4981-b2b8-e4323a7d785b · outbound

This paper cites N24news: A new dataset for multimodal news classification.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs N24news: A new dataset for multimodal news classification

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:57.854802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:57.854802Z digest=sha256:b0eb4efa62278521187a3fb136b7dd82d77fc684c70acc929cac60da76104ca4

Observation 0fb1d256-49c1-4dae-a81c-7c91b7fb32d9 · outbound

This paper cites Uniir: Training and benchmarking universal multimodal information retrievers.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Uniir: Training and benchmarking universal multimodal information retrievers

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:57.950522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:57.950522Z digest=sha256:4553b319cbca217f1015614cbc19b552547b921854aa5895d105e0d53f72593c

Observation 25b135c0-e703-4dfe-ae65-1a345f848257 · outbound

This paper cites Fashion iq: A new dataset towards retrieving images by natural language feedback.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Fashion iq: A new dataset towards retrieving images by natural language feedback

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:58.024964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:58.024964Z digest=sha256:7df3290705b7779d61c854ae44873239a0e96003f2503b99580a4791165b8064

Observation 95610b8e-fecd-4fb0-88d1-7d010b0017b9 · outbound

This paper cites Sun database: Large- scale scene recognition from abbey to zoo.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Sun database: Large- scale scene recognition from abbey to zoo

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:58.105500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:58.105500Z digest=sha256:8e2f4a57b1df0cb2c890a25b37b5eb955f46c62e03d50b51a4cd9cae65893b5e

Observation e645e263-3f3f-4f85-b91b-a49573b0af2c · outbound

This paper cites Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:58.191649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:58.191649Z digest=sha256:097745942345693f8d921be8216599b49517bf13de38b4d109dd79c0c58f21e3

Observation 63a04c57-ea50-4ac5-8c7b-f62b7f439f89 · outbound

This paper cites Topv: Compatible token pruning with inference time optimization for fast and low-memory multimodal vision language model.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Topv: Compatible token pruning with inference time optimization for fast and low-memory multimodal vision language model

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:58.293990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:58.293990Z digest=sha256:bcb9c244e6a9a2283f5d7de3ff376de2264bc247f137bad60dfd76ccc4cf05e7

Observation 0ab8ebf9-b810-4d6a-8d88-74517a54428a · outbound

This paper cites an unresolved cited work.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Unresolved cited work

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:58.375026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:58.375026Z digest=sha256:1e18d11112cb0506e15422c4266b4f61df25b61732409ff28bd48b710716033d

Observation 8455edb3-1e95-443d-b124-2f249ab84a32 · outbound

This paper cites Visrag: Vision-based retrieval-augmented generation on multi- modality documents.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Visrag: Vision-based retrieval-augmented generation on multi- modality documents

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:58.475413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:58.475413Z digest=sha256:bf69fa1b1e4e448c3b12b5b321d7f419c1dbf80320d13799308558bc78f44060

Observation 57d57324-eb6c-4f45-af97-3a5a3dadc5b7 · outbound

This paper cites RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:58.564725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:58.564725Z digest=sha256:9b0b3285414186c168a13306fe8908d1c0bb018b7f311065ccf6d9b0cb87507d

Observation 1209586c-cb5d-475d-a83e-68d6e4e3ea33 · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv preprint arXiv:2405.17220, 2024.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv preprint arXiv:2405.17220, 2024

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:58.673363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:58.673363Z digest=sha256:de50194c7655679bf5e41ca78d1670cdeb1201613ff98f427c4517817474c880

Observation ba776f33-078c-43d5-a1df-22b946cefd29 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:58.781116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:58.781116Z digest=sha256:ad0e6d0c790e65c3640e87aaed510fe60cd4ad4a350f2c9bebaad0ff7c7503da

Observation 96607e15-d8cb-43f3-a865-9623378023ab · outbound

This paper cites Sigmoid loss for language image pre- training.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Sigmoid loss for language image pre- training

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:58.884879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:58.884879Z digest=sha256:581666035227aff72789463cfb6d4fa47d5a7bdfb019461ae25b0da462fb2490

Observation 3b079c7f-896a-4618-9d41-70ade8542c72 · outbound

This paper cites Long-clip: Unlocking the long-text capability of clip.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Long-clip: Unlocking the long-text capability of clip

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:58.969922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:58.969922Z digest=sha256:614cbfbf238044408ee25796911bbf9f8cc479d84c9354a8bf18dd71b8da7133

Observation 54fab929-540e-4984-9a36-29b9b97611d0 · outbound

This paper cites Notellm-2: Multimodal large representation models for recommendation.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Notellm-2: Multimodal large representation models for recommendation

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:59.051193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:59.051193Z digest=sha256:71ad00931c6def87c9caedb49c369335775e90767d0bd64e97aeb341d89a5462

Observation 64dd915b-aee1-431a-aa44-09440a33f759 · outbound

This paper cites MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:59.152618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:59.152618Z digest=sha256:28467ca80431c45368f06a3279c0affd08d2d3b39193fc0596f675f9876b8c87

Observation ba33d555-e37f-4f04-b8c6-f4f9b1180065 · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:59.229137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:59.229137Z digest=sha256:1ca02fa98d2b345f4dc3403d92782b17166d16fc64e8c5ea3df27953957135f8

Pith citing papers

No inbound Pith citation observations are available.