Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:31:16.361763Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 4 inbound Pith citation observations for arXiv:2501.04670.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:31:16.361763Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:37:50.658662Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T19:16:08.251671Z
100 of 111 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 90321269-2cf5-45dd-8c30-29769e0f55c9 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57098b75-34aa-4cda-b4e7-7b1cd06be040 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs PaLM 2 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6732e9ff-f799-477a-b39e-b17d87135789 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Burst: A benchmark for unifying object recognition, seg- mentation and tracking in video
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 527fb9b8-b1c8-4a4e-82b0-f3680f4ab722 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Qwen Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f0b80c-7712-437f-a486-824eb1d13354 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Reliable feature matching across widely separated views
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbb8b745-e400-45a6-94c5-659900ae7d70 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c640174-16b3-4829-8fc4-693341dedb1f · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs The wildtrack multi-camera person dataset
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b9a60bd-82cc-4102-8d1a-4d9ae8efb204 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40c53023-fc59-47b8-b972-1b6200e056da · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa1378a-d473-428b-8b1f-26c93e7e4cd1 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Improved Baselines with Momentum Contrastive Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 087cd7be-d6e1-4c05-b617-a14ff3b46f4a · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98aaa744-62e6-45b3-ba3d-348941031cf1 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82313edb-d6a3-40f6-b6e5-7184ac430632 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Xtuner: A toolkit for efficiently fine-tuning llm
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9d028b7-1faa-4f7d-a738-ad5e6cfcf9f3 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs SportsMOT: A Large Multi-Object Tracking Dataset in Multiple Sports Scenes
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b795c884-64ee-4230-9b90-17439a52ab77 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aafa7cf-3f03-411d-80f3-bdd9d80f9245 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Mose: A new dataset for video object segmentation in complex scenes
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2a94e46-8837-4505-ab0f-dd70501d4c16 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Lasot: A high-quality benchmark for large-scale single ob- ject tracking
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b78bebe6-111c-4bc4-96c5-e946d1114c2e · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Qd- track: Quasi-dense similarity learning for appearance-only multiple object tracking
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 991cd2d1-5a1d-44c3-bd66-a47b2a72408a · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed753e0f-cf65-4505-83dc-470f8ffcd8d8 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 998dafa7-9ee5-4994-a3f6-249f9c42bce6 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Blink: Multimodal large language models can see but not perceive
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af339474-5a8b-4b80-b5bb-cd653ae9e5e7 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Stere- oscan: Dense 3d reconstruction in real-time
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68f9d8b3-3e6b-4d6d-b7f6-c5f60966aeec · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Making the V in VQA matter: El- evating the role of image understanding in Visual Question Answering
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08afac73-5868-419c-91d3-ab286e708614 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Momentum contrast for unsupervised visual rep- resentation learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c962e9bc-be23-421b-a005-27e52580685c · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs 3d-llm: Inject- ing the 3d world into large language models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 655bffd6-4962-4a33-b1d2-cfaaea2ab7e4 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Multi- view detection with feature perspective transformation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f294d1d9-4e9d-4981-af11-cfb6c427cd19 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs LoRA: Low-Rank Adaptation of Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c70a6cd-802c-4b89-891a-bbdc4fce25c7 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Global instance tracking: Locating target more like hu- mans
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b8e7c75-b2ca-4604-b241-27ee36627fd3 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 296f46c4-2cf7-4626-894a-f6ffc0a7ef99 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Segment Anything
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab6b5c80-ea47-4e38-ab44-fadae128fc91 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs LISA: Reasoning Segmentation via Large Language Model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c41efef0-c3eb-49d1-978d-7ad836b74b1e · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Obelics: An open web-scale filtered dataset of interleaved image-text documents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f37bc49-f74a-463b-9316-54397e374662 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19835f3a-e59f-45cf-b729-4f0f20fd8f72 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7f28594-1b24-47e6-a3d8-0760c86a7a36 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8ad67d0-7913-4830-9102-da39db3f9fd5 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Matching anything by segmenting anything
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3226db4b-1b34-46be-b75e-c589c6d76390 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Video k-net: A simple, strong, and unified baseline for video segmentation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ad672a7-db4e-4e7a-a66d-bd1b09229616 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Tube-link: A flexible cross tube framework for universal video seg- mentation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba73303b-7d66-415b-a8be-687f776215c7 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Omg-seg: Is one model good enough for all segmen- tation? In CVPR, 2024
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93be29d1-cf5f-4e59-bd35-17fcb218fd6f · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Evaluating Object Hallucination in Large Vision-Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6252404-e2de-48ef-9d72-f74273b99ac4 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c079c19d-878a-425c-87cd-a28d0ca3fe4e · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Llama-vid: An image is worth 2 tokens in large language models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40e096ad-5d63-4d26-8431-b37d1d41f4d0 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Mon- key: Image resolution and text label are important things for large multi-modal models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fbc29ad-f07b-41d6-bf44-229c0d1ae69e · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 337fc6be-176a-4f92-b306-3550a81e47ae · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Vila: On pre-training for vi- sual language models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09028879-9b50-4a91-9990-f96ac8d269d4 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs MM-VID: Advancing Video Understanding with GPT-4V(ision)
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61998029-218c-4659-9462-e1c4b65bd4bd · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 292a7dbe-79ce-45b2-baa7-06880edaf48b · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Improved baselines with visual instruction tuning, 2023
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 081265e9-c841-44e9-b891-b5c739e4c192 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Visual Instruction Tuning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a912c0c-492a-4a84-9d69-346f68122dbe · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Improved Baselines with Visual Instruction Tuning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bef9819-8ab3-4451-8d33-f350124ceb12 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f675420-c447-4db8-ae8e-8ca49c964033 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Mmbench: Is your multi- modal model an all-around player? In ECCV, 2024
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5c77256-dca2-4127-bd1b-abe53040b7c9 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfb87a1a-2332-4769-8d31-8c1dc71dc8c7 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b2450fe-3d50-4125-b613-72a606a21916 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ebe05e8-b82c-4828-ae81-e366e7491c40 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7748671-a59c-48c5-8ab4-38f160ab2934 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Large-scale video panoptic segmentation in the wild: A benchmark
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82bceed4-766d-4f3b-b424-12ac6e0820a8 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs MOT16: A Benchmark for Multi-Object Tracking
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bce8364c-c7ea-4525-8f9b-551d991afb5f · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs GPT-4 Technical Report
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e445485-bdbf-41f9-809f-4d303609d25e · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs DINOv2: Learning Robust Visual Features without Supervision
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f704c7-6405-4818-9e84-6f40d598c3c7 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs VastTrack: Vast Category Visual Object Tracking
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18d55902-a0ff-41ff-a7a7-43f10da9bb6e · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Occluded video instance segmentation: A benchmark
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb51cf99-ca08-4d5d-85d9-fa50e9f7849b · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Chatvtg: Video temporal grounding via chat with video dialogue large language models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d706e75-afce-498b-ab60-d015d56c5b6b · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Learn- ing transferable visual models from natural language super- vision
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a038d13-666f-4f43-9206-827a491c6539 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Am-radio: Agglomerative vision foundation model reduce all domains into one
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4dd20f79-1bd9-40c9-86b7-ff622992150e · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Glamm: Pixel grounding large multimodal model
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation acaa19dc-91b8-4b4c-8d6f-7a9a736a1174 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs SAM 2: Segment Anything in Images and Videos
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8093b0b-6d46-4454-a930-67eba09d1e5f · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15e62f27-e382-42df-a63e-5c2e58abe1f4 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Towards vqa models that can read
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2e04ca02-d7c2-485d-ba2e-3d3bc77af7c3 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Moviechat: From dense token to sparse memory for long video understanding
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d508a717-b65a-4bbb-b85a-e51aa83eb981 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language Model
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 28d53b87-c524-41fd-847e-9e67d4a81f1f · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 741d05d3-ed3a-43a4-b5c6-86c6c6eef009 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Internlm: A multilingual language model with progressively enhanced capabilities, 2023
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 70bfd6b4-b1e2-45ab-8335-0f85daad112a · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29a3ebcb-727b-4ba6-a42a-a9f142c12e59 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d1038f15-88fa-4e47-91c3-82bf4777c9d0 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs LLaMA: Open and Efficient Foundation Language Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff64e833-3947-4420-9219-c3f1ecca0f15 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cddff9e9-2f88-41a1-88f0-c03453586124 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Long-term tracking in the wild: A benchmark
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d42da109-7df7-40f7-bff1-cb300db98040 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10237edd-c20c-49b6-8230-70466c8f5a45 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Ov-vis: Open-vocabulary video instance seg- mentation
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2b3e831b-1d01-41c5-82b8-b4eb10e74db0 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1767ed95-8310-40fd-8c48-e2186ab2746a · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f335d46-a268-4bc5-9f39-8e83f6b3e7bc · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Chain-of-thought prompting elicits reasoning in large lan- guage models
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d407407d-f374-4227-96b1-de8c707083f1 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Con- vnext v2: Co-designing and scaling convnets with masked autoencoders
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8f71c8b3-5499-45a5-afa6-0eb30d4763b3 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs In defense of online models for video in- stance segmentation
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d33eb56b-533a-47a9-ad21-f925e4066063 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 16d49c21-66b4-4731-ba08-7972d6572549 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Pointllm: Empowering large language models to understand point clouds
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation db653ac8-98b8-4151-a588-ddc956c61c2f · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs xgen-mm (blip-3): A family of open large multimodal models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92d8533a-d74d-4fc6-9d86-fafc7586d619 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Towards grand unification of object tracking
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 66396e52-a2d5-4f24-a4e7-6a73965e386f · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Universal instance perception as object discovery and retrieval
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation aa0dda1e-ef11-48f1-93d8-57b02bde466b · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Qwen2 Technical Report
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4865195-a39c-45e2-8cd1-a52f834ee2c4 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Video instance segmentation
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fb206a06-e5a0-4e86-87c7-d27abfecece4 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc93b8b2-4ade-4062-ad9d-44bd2c7f7811 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c73889c-c16d-4e12-81cd-b7ad9e9f7d66 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Ctvis: Consistent train- ing for online video instance segmentation
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 544897a7-4d45-4c5d-8b9d-c1f9c41dced9 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Yi: Open Foundation Models by 01.AI
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e86708ae-6156-472b-a81c-37c15e0a4080 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Bdd100k: A diverse driving dataset for heteroge- neous multitask learning
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 702b7792-0929-453e-b6a4-259a6ce811b7 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ca3cf87-e249-4e75-85ff-31c580730775 · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Osprey: Pixel understanding with visual instruction tuning
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c8d6adaa-7b91-48ae-b276-bb9dd7ce884c · outbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1c18b3a-35d8-48a1-9128-c02701f41d19 · inbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd2f3850-3373-487a-b4ea-3993548ed273 · inbound
Dense360: Dense Understanding from Omnidirectional Panoramas Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cab53f23-af9c-4c36-83a5-25d31c6b7cc9 · inbound
MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8a408e2-9f44-4a31-952e-edf973f705e3 · inbound
CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.