Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T04:20:59.229137Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 0 inbound Pith citation observations for arXiv:2602.05275.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T04:20:59.229137Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 109 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 166a2c47-85ae-4c04-906c-9eee29fe86a9 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be59d751-9b27-45f1-93fc-590044a5f705 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Coig-cqia: Quality is all you need for chinese instruction fine-tuning, 2024
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05d1e918-7453-4f08-aeda-dba483f41331 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models.Advances in neural information processing systems, 32, 2019
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d3a293f-257b-4725-9561-cd6f6ba388f5 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Baai-mtp dataset
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dccbc202-eb71-431c-a90a-e929ada7b54e · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b69c64be-4bfc-4880-8c75-5a0d9ff684d0 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Honeybee: Locality-enhanced projector for multimodal llm
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7868e2cf-7529-46d5-b645-eb23bab88664 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Webqa: Multihop and multimodal qa
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c243006-424a-4db2-b9c6-014a5ba99d4d · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2335a39d-00c1-49f2-ada7-3d9eccb47a02 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Sharegpt4v: Improving large multi-modal models with better captions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acab2e52-707d-4483-a256-c636ab96fd9b · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Reproducible scaling laws for contrastive language-image learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ce5ec27-ebf6-4054-8c4b-ab3e8b33442d · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Think then embed: Generative context improves multimodal embedding.arXiv preprint arXiv:2510.05014, 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24da3b8c-1f04-49fd-abd6-26b21d0ed7f5 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Visual dialog
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3afd858a-3cbb-45c8-9c88-1b34042db5d1 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Imagenet: A large-scale hierarchical image database
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48688fd7-db64-449c-a223-2b57973041c5 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Bert: Pre-training of deep bidi- rectional transformers for language understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 179adbdc-b9d6-47a0-a487-8597ebc8b8b6 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Pact: Pruning and clustering-based token reduction for faster visual language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31325f37-7599-413e-91ee-3397c878c3b1 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs An image is worth 16x16 words: Transformers for image recognition at scale
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43981ea6-2a2e-434c-a555-09f7d8b36f02 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs The pascal visual object classes challenge: A retrospective.International journal of computer vision, 111(1):98–136, 2015
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46bf6137-d634-43fd-a7ac-5214bd75eebb · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs ColPali: Efficient Document Retrieval with Vision Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 183a65f1-5665-498b-9cbf-7a05848fb385 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36da3974-7e5d-4c11-8b8a-6abe816d0e2d · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Vl-clip: Enhancing multimodal recommendations via visual grounding and llm-augmented clip embeddings
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae888698-4138-4df4-bfa0-ca4c4b4bc9d8 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d04b80f-cadd-4690-bdf0-76de269deeb4 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Breaking the modality barrier: Universal embedding learning with multimodal llms
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09dd4733-71ab-4e64-9960-30af4ee00814 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Unime-v2: Mllm-as-a-judge for universal multimodal embedding learning.AAAI, 2026
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa8ff887-223a-4a26-965b-f40dcfcec6f0 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Vizwiz grand challenge: Answering visual questions from blind people
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82da0299-3d75-4c2d-ab49-b347be8f59b6 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Efficient Multimodal Learning from Data-centric Perspective
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96a5728c-d82b-42ed-b2c6-63046e03b7c3 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs The many faces of robustness: A critical analysis of out-of- distribution generalization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d99ed3e6-fe96-4b1b-8de8-809fbf6fa007 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Natural adversarial examples
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d802671-6139-46ae-a30b-c63e2d2b8bd8 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36:31096–31116, 2023
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91081082-ff54-49d8-b4d1-911425dcc288 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Open-domain visual entity recognition: Towards recognizing millions of wikipedia entities
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfb612ed-935a-4f4b-80d7-b1570230c04f · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bc95e34-6faa-4aaf-8c14-fbb97c7c3366 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs VideoRAG: Retrieval-Augmented Generation over Video Corpus
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acb3f71a-b666-48ae-aaf1-634564865cbe · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Scaling up visual and vision-language representation learning with noisy text supervision
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 977c4326-5f16-48ca-b8b6-6358a3d87dcf · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Rzenembed: Towards comprehensive multimodal retrieval.CoRR, abs/2510.27350, 2025
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9095c6f-3ea3-4a98-b77b-b176b4c5dc82 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs E5-V: Universal Embeddings with Multimodal Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12f5a73a-ca73-436e-acd4-ea478d6def41 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Vlm2vec: Training vision-language models for massive multimodal embedding tasks.ICLR, 2025
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25b78c6e-7297-431e-842c-9149d7217160 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Referitgame: Referring to objects in photographs of natural scenes
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd25e9aa-bbbf-412f-b772-e53181d135bc · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs The hateful memes challenge: Detecting hate speech in multimodal memes.Advances in neural information processing systems, 33:2611–2624, 2020
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8b59236-ee56-45db-b021-deaa1d229b16 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ba3d8cc-e5eb-4ab6-b1c9-027339cbd6c6 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 213c22ef-aa82-46dd-97b9-45934dc05b11 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Gonzalez, Hao Zhang, and Ion Stoica
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c4ddbba-ec11-4314-9fdd-3a0ff0b81ff3 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Llave: Large language and vision embedding models with hardness-weighted contrastive learning.CoRR, abs/2503.04812, 2025
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 929d70e2-7b2a-48cc-a0c1-b4a09760ae73 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Ume-r1: Exploring reasoning-driven generative multimodal embeddings.arXiv preprint arXiv:2511.00405, 2025
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1a1b7cc-4327-4d06-8e40-7aaa6efa2abc · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Building and better understanding vision-language models: insights and future directions
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8737e8ec-8fc9-461a-97ec-70b28ff7d280 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28e058cf-59cc-429f-a3e6-6b3e1948a423 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b1f1732-8e9b-4c7c-97ae-e17119a92bef · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f1bd9c3-28b3-4f0a-9308-6f87286f58e3 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcfcfc53-0e48-4a26-963a-ef15269b830b · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Silkie: Preference Distillation for Large Visual Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc433143-8e98-4778-8378-95c28f42aceb · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f805cdb2-ece6-4b0e-ad8e-6766ace1d552 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Monkey: Image resolution and text label are important things for large multi-modal models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd0e4e57-7762-4cbd-815f-a76f08928956 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c2cb28b-4ad8-41a6-9bed-55a112eca2a0 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Microsoft coco: Common objects in context
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9143fca5-9bfc-428a-a83b-e062fc6b8bf0 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Gres: Generalized referring expression segmentation
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 190b143f-5bb8-4e52-b0fd-96324161291e · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2de17b5a-c7b5-472b-9aba-918622fe9e6f · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Visual news: Benchmark and challenges in news image captioning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12f91ffe-ddfb-400a-bc06-bee301d4359a · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Improved baselines with visual instruction tuning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f072e3e-75bf-43f4-a3e2-bcaf95f242df · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Exploiting transformation invariance and equivariance for self-supervised sound localisation
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2c6451b-e443-4568-9f64-85f32a287c39 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Edis: Entity-driven image search over multimodal web content
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00fcc7cc-c367-458a-93c4-218c97b71f19 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Lamra: Large multimodal model as your advanced retrieval assistant.CVPR, 2024
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a0d4ea-cc76-44ac-9e27-733a7b3bd8cb · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Image retrieval on real-life images with pre-trained vision-and-language models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a45aa92c-f61c-4a21-931f-c62c279a5b61 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58ad22d6-b653-4236-9404-83176f58b5ae · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Unifying Multimodal Retrieval via Document Screenshot Embedding
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b18985b3-4092-4a88-8509-22269840a629 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Visa: Retrieval augmented generation with visual source attribution
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86b9a751-c795-4086-9ac3-21f8e45aeff7 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Mmlongbench-doc: Benchmarking long-context document understanding with visualizations.Advances in Neural Information Processing Systems, 37:95963–96010, 2024
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c34fcdb9-e9d0-496e-b7d3-556391baf8b7 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Vidore benchmark v2: Raising the bar for visual retrieval.arXiv preprint arXiv:2505.17166, 2025
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7064aa75-1c13-4bb3-b89f-41d8c54ce929 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Generation and comprehension of unambiguous object descriptions
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8adf7aa-019e-4dbc-b2ac-7fc6aa70587c · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3f7e06e-2efe-48a1-8021-d91f584d35a6 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74284f40-3807-4964-9064-ede3c1dffadd · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs In- fographicvqa
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0468aa6e-2419-4b77-92df-d4c586a2df12 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Docvqa: A dataset for vqa on document images
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b9ed98a-34ad-43c7-a485-f8ac7fb784f7 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07739cff-4fdd-40ed-a91e-678f06fad60a · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Ops-mm-embedding-v1
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b96811c7-27b7-4610-86cc-2b0c40849d8a · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c18fb213-721c-4e00-80c4-94d9f798838a · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Learning transferable visual models from natural language supervision
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d885533-b961-4aff-ac6c-87c1f239c3c9 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs A- okvqa: A benchmark for visual question answering using world knowledge
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 004919ca-0b3c-4749-9a9e-1e8998d48a18 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Objects365: A large-scale, high-quality dataset for object detection
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aafa2c93-40bf-4735-a64c-b8f7d18dbf5d · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Sharegpt-chinese-english-90k: A bilingual chinese-english human-machine dialogue dataset
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e76efb8-36ed-4ea1-93e3-bfc2a5ee58ce · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Towards vqa models that can read
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24b1a7b5-fabe-48c9-99c0-5ccc3d4e1fe9 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faabcb24-c86a-452e-a640-9a27f8a5cd1d · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Breaking the batch barrier (b3) of contrastive learning via smart batch mining
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fe26576-c358-45b5-a0c1-e3d8593f75b0 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Representation Learning with Contrastive Predictive Coding
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 596c61a7-975b-44b2-82cc-8b4996b6aef3 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs V3det: Vast vocabulary visual detection dataset
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 260f681b-8edd-4671-8ead-82b65f1cc844 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abafe370-d164-4aaa-a5ea-aae9a7ffed09 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b61da42-886d-4981-b2b8-e4323a7d785b · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs N24news: A new dataset for multimodal news classification
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fb1d256-49c1-4dae-a81c-7c91b7fb32d9 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Uniir: Training and benchmarking universal multimodal information retrievers
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25b135c0-e703-4dfe-ae65-1a345f848257 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Fashion iq: A new dataset towards retrieving images by natural language feedback
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95610b8e-fecd-4fb0-88d1-7d010b0017b9 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Sun database: Large- scale scene recognition from abbey to zoo
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e645e263-3f3f-4f85-b91b-a49573b0af2c · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63a04c57-ea50-4ac5-8c7b-f62b7f439f89 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Topv: Compatible token pruning with inference time optimization for fast and low-memory multimodal vision language model
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ab8ebf9-b810-4d6a-8d88-74517a54428a · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Unresolved cited work
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8455edb3-1e95-443d-b124-2f249ab84a32 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Visrag: Vision-based retrieval-augmented generation on multi- modality documents
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57d57324-eb6c-4f45-af97-3a5a3dadc5b7 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1209586c-cb5d-475d-a83e-68d6e4e3ea33 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv preprint arXiv:2405.17220, 2024
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba776f33-078c-43d5-a1df-22b946cefd29 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96607e15-d8cb-43f3-a865-9623378023ab · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Sigmoid loss for language image pre- training
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b079c7f-896a-4618-9d41-70ade8542c72 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Long-clip: Unlocking the long-text capability of clip
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54fab929-540e-4984-9a36-29b9b97611d0 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Notellm-2: Multimodal large representation models for recommendation
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64dd915b-aee1-431a-aa44-09440a33f759 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba33d555-e37f-4f04-b8c6-f4f9b1180065 · outbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.