Pith. sign in

Paper Citation Record · LEDGER

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

As of 21 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 4 inbound Pith citation observations for arXiv:2501.04670.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04670 v3

Coverage vector

measured 100 of 111 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:31:16.361763Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:37:50.658662Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T19:16:08.251671Z

Reference resolution

100 of 111 outbound references displayed

  • verified exact3
  • verified fuzzy18
  • unresolved79
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 90321269-2cf5-45dd-8c30-29769e0f55c9 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.840634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.840634Z digest=sha256:e77a90a701043a44c209096fa5519b952c01c2ffee71962e690ec7d6207df9a4

Observation 57098b75-34aa-4cda-b4e7-7b1cd06be040 · outbound

This paper cites PaLM 2 Technical Report.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs PaLM 2 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.844794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.844794Z digest=sha256:c6f557359241428de38ba42b4c6447fc14296ffa95d19d07b4655ed683984398

Observation 6732e9ff-f799-477a-b39e-b17d87135789 · outbound

This paper cites Burst: A benchmark for unifying object recognition, seg- mentation and tracking in video.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Burst: A benchmark for unifying object recognition, seg- mentation and tracking in video

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.848867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.848867Z digest=sha256:53b5fbb39431d2d656f87a0415292f1598f9189b4b2138f3318bd8251e641dd1

Observation 527fb9b8-b1c8-4a4e-82b0-f3680f4ab722 · outbound

This paper cites Qwen Technical Report.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.852616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.852616Z digest=sha256:44de16afa74e904043dbe6a2dc260f9569e49c4e2b2b049bda39315efceee7ff

Observation 06f0b80c-7712-437f-a486-824eb1d13354 · outbound

This paper cites Reliable feature matching across widely separated views.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Reliable feature matching across widely separated views

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.856343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.856343Z digest=sha256:51458f86df57e108ea9c22d9ce2b2cc1b0574c6ca52038e75609ffb77e81bfdd

Observation fbb8b745-e400-45a6-94c5-659900ae7d70 · outbound

This paper cites Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.860053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.860053Z digest=sha256:4de48b18e10939a842dd30e153e85653a0c31cdd55e132c26ca58eeb6a06f656

Observation 7c640174-16b3-4829-8fc4-693341dedb1f · outbound

This paper cites The wildtrack multi-camera person dataset.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs The wildtrack multi-camera person dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.863608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.863608Z digest=sha256:c3343f8ac9e46dccb83cf22f8b12f874f74e4e33a1f7297e04f4dc4e50980dbe

Observation 1b9a60bd-82cc-4102-8d1a-4d9ae8efb204 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.867604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.867604Z digest=sha256:d8d3536382455f2c398be5d76c805c4747e0f47e89e4a84540181fe059a29747

Observation 40c53023-fc59-47b8-b972-1b6200e056da · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.871287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.871287Z digest=sha256:5b7cb0e46b5fbf969ac6b79ad59f45e2394ad09376d0b53cd555d019111b7182

Observation 2fa1378a-d473-428b-8b1f-26c93e7e4cd1 · outbound

This paper cites Improved Baselines with Momentum Contrastive Learning.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Improved Baselines with Momentum Contrastive Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.874974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.874974Z digest=sha256:81464d30ab11a22fee34412326ed217f0c60696f3db7a60e07c0eca24d7572bb

Observation 087cd7be-d6e1-4c05-b617-a14ff3b46f4a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.878718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.878718Z digest=sha256:1f33526c565523b823a468cffc367135f167effc37e3d80686cff08e37b334ec

Observation 98aaa744-62e6-45b3-ba3d-348941031cf1 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.882669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.882669Z digest=sha256:efe83358bd6f7c76f13fc54fc21375b016895a24f51a842e356bb2a51735dbaa

Observation 82313edb-d6a3-40f6-b6e5-7184ac430632 · outbound

This paper cites Xtuner: A toolkit for efficiently fine-tuning llm.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Xtuner: A toolkit for efficiently fine-tuning llm

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.886244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.886244Z digest=sha256:53824323d29cee9a7537c49338b8f61fdc706c6e62d143ac7888eb2f0888f3de

Observation f9d028b7-1faa-4f7d-a738-ad5e6cfcf9f3 · outbound

This paper cites SportsMOT: A Large Multi-Object Tracking Dataset in Multiple Sports Scenes.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs SportsMOT: A Large Multi-Object Tracking Dataset in Multiple Sports Scenes

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:31:17.057427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:15.889562Z digest=sha256:22477d6b0ce8833dbd5daa39997b1dfd3654eabfa88bd503afdaf567d9ef17ff

Observation b795c884-64ee-4230-9b90-17439a52ab77 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.893065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.893065Z digest=sha256:2fd015f163bdea60375114f88d7f3ee15e93a402197093e0711abf0816629dca

Observation 4aafa7cf-3f03-411d-80f3-bdd9d80f9245 · outbound

This paper cites Mose: A new dataset for video object segmentation in complex scenes.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Mose: A new dataset for video object segmentation in complex scenes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.896518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.896518Z digest=sha256:77f0c7bfea6ba72d5050ddc8834b4c4ef22ab6d1027673e69860eef50c42450b

Observation b2a94e46-8837-4505-ab0f-dd70501d4c16 · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single ob- ject tracking.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Lasot: A high-quality benchmark for large-scale single ob- ject tracking

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.900444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.900444Z digest=sha256:e0ea636a977353e1abb4963d8e7a3b2ab28b156a60ff4fd661bdb86769c8f21f

Observation b78bebe6-111c-4bc4-96c5-e946d1114c2e · outbound

This paper cites Qd- track: Quasi-dense similarity learning for appearance-only multiple object tracking.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Qd- track: Quasi-dense similarity learning for appearance-only multiple object tracking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.904698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.904698Z digest=sha256:f5b2cf79290aca6ed9df7c85e4be39028dfaebe63b66bf9ff26ef0e21c28d755

Observation 991cd2d1-5a1d-44c3-bd66-a47b2a72408a · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.908538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.908538Z digest=sha256:15b414909eef77cfd27ac644c8aaf7118322858749b5cae078c43aa9373ca0a0

Observation ed753e0f-cf65-4505-83dc-470f8ffcd8d8 · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.912278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.912278Z digest=sha256:ab1a3663282ef69ae09bb578ac423be311593f7dae5b57dcfdf99b23f83f6e9a

Observation 998dafa7-9ee5-4994-a3f6-249f9c42bce6 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Blink: Multimodal large language models can see but not perceive

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.915924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.915924Z digest=sha256:18afb1443d2fd2882743b908d02ca342c5f4f95ebfb15bb33a04bada9716f8ec

Observation af339474-5a8b-4b80-b5bb-cd653ae9e5e7 · outbound

This paper cites Stere- oscan: Dense 3d reconstruction in real-time.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Stere- oscan: Dense 3d reconstruction in real-time

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.918863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.918863Z digest=sha256:4ddd7781db89d20b9ad906ddad93279a2ab96de6a06a686b362d5eb26120966e

Observation 68f9d8b3-3e6b-4d6d-b7f6-c5f60966aeec · outbound

This paper cites Making the V in VQA matter: El- evating the role of image understanding in Visual Question Answering.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Making the V in VQA matter: El- evating the role of image understanding in Visual Question Answering

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.921944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.921944Z digest=sha256:4146a718c6df85c425a5a194b38c7c9af6ad05da977f146ed631da64d1614a1f

Observation 08afac73-5868-419c-91d3-ab286e708614 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Momentum contrast for unsupervised visual rep- resentation learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.924741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.924741Z digest=sha256:a7afcdffefa6e448e7a2e52da3605f7724784f1138f6b6b57613b9bdb4020c44

Observation c962e9bc-be23-421b-a005-27e52580685c · outbound

This paper cites 3d-llm: Inject- ing the 3d world into large language models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs 3d-llm: Inject- ing the 3d world into large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.927784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.927784Z digest=sha256:d2d5779105f2b1c03369c69474488011669a442aa82bad1fceb707b4bc933304

Observation 655bffd6-4962-4a33-b1d2-cfaaea2ab7e4 · outbound

This paper cites Multi- view detection with feature perspective transformation.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Multi- view detection with feature perspective transformation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.931189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.931189Z digest=sha256:f93dcc3b6be1b57a059b0f04d5658d3fb0eeaf24c84388d16b6349731f6b3827

Observation f294d1d9-4e9d-4981-af11-cfb6c427cd19 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs LoRA: Low-Rank Adaptation of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.934370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.934370Z digest=sha256:e0e6cb7725ac93e879039c4f96bc0985211252d8405a6a36893ad3a016a39b49

Observation 3c70a6cd-802c-4b89-891a-bbdc4fce25c7 · outbound

This paper cites Global instance tracking: Locating target more like hu- mans.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Global instance tracking: Locating target more like hu- mans

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.937479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.937479Z digest=sha256:b161bc181fbde8eb3eb1658de5cedfddf02fe70731a8d424305cd2645d67654f

Observation 1b8e7c75-b2ca-4604-b241-27ee36627fd3 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.941044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.941044Z digest=sha256:d8f7537213a4eb3b8c2979f4cf624e92de52409322be88d9250d37c41b8cd4c6

Observation 296f46c4-2cf7-4626-894a-f6ffc0a7ef99 · outbound

This paper cites Segment Anything.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Segment Anything

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.947653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.947653Z digest=sha256:993f0228834f29864a7a6e0a5bc8b24c50a4f4605ede7e05008a661b837b00c3

Observation ab6b5c80-ea47-4e38-ab44-fadae128fc91 · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs LISA: Reasoning Segmentation via Large Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.951403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.951403Z digest=sha256:e434343f48665f08507c57d86c5fd413c39d00f25982374f5662ef31262934ad

Observation c41efef0-c3eb-49d1-978d-7ad836b74b1e · outbound

This paper cites Obelics: An open web-scale filtered dataset of interleaved image-text documents.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Obelics: An open web-scale filtered dataset of interleaved image-text documents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.955297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.955297Z digest=sha256:72228baab08f4e27648d58c964b78165845bf4fc3b45e0aafdfbe28a549d4079

Observation 6f37bc49-f74a-463b-9316-54397e374662 · outbound

This paper cites NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.958685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.958685Z digest=sha256:d39013ba9120d3a95b1a72e00c56b1a1b4041be2df7cc3fd87626c6ffe30fb5c

Observation 19835f3a-e59f-45cf-b729-4f0f20fd8f72 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.962372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.962372Z digest=sha256:e06a46265253b77e98a8f959f387f96aa20a833d64cb414db6d231394611f910

Observation f7f28594-1b24-47e6-a3d8-0760c86a7a36 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.965901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.965901Z digest=sha256:92893270ada126a7096003d1a85012627628189b924713d427d64969f15e0b54

Observation a8ad67d0-7913-4830-9102-da39db3f9fd5 · outbound

This paper cites Matching anything by segmenting anything.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Matching anything by segmenting anything

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.969478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.969478Z digest=sha256:be51a64c6b7e1fb9e6bda6997ee65b3b134ee8f91690748b48adbeeeabf0551e

Observation 3226db4b-1b34-46be-b75e-c589c6d76390 · outbound

This paper cites Video k-net: A simple, strong, and unified baseline for video segmentation.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Video k-net: A simple, strong, and unified baseline for video segmentation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.973849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.973849Z digest=sha256:35549082292114b2873132aae99e5d4dc6af6fb67729742b6bd9643d75e2517f

Observation 4ad672a7-db4e-4e7a-a66d-bd1b09229616 · outbound

This paper cites Tube-link: A flexible cross tube framework for universal video seg- mentation.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Tube-link: A flexible cross tube framework for universal video seg- mentation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.977541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.977541Z digest=sha256:7e8bf50775fbb0fb38c8b3c517caf7bbd0c6a835333bed1d1767045ae9b7c9c1

Observation ba73303b-7d66-415b-a8be-687f776215c7 · outbound

This paper cites Omg-seg: Is one model good enough for all segmen- tation? In CVPR, 2024.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Omg-seg: Is one model good enough for all segmen- tation? In CVPR, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.981162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.981162Z digest=sha256:4279a01ad597eb5046e3b02152ef2c7015a478a8c8737939569c1ba90b836b45

Observation 93be29d1-cf5f-4e59-bd35-17fcb218fd6f · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Evaluating Object Hallucination in Large Vision-Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.984902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.984902Z digest=sha256:512b655e2edc3d7560ffed01f486d81c84a98cd7f8ad2d095374f56266e72334

Observation e6252404-e2de-48ef-9d72-f74273b99ac4 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.988827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.988827Z digest=sha256:d287d33c77f684649174b41ec8ff1dfdb4824716439e1efba446032b90ede7e3

Observation c079c19d-878a-425c-87cd-a28d0ca3fe4e · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Llama-vid: An image is worth 2 tokens in large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.993094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.993094Z digest=sha256:355aef97cb17110c550a210947c9b88203d4298b065e934092802f7f816904e9

Observation 40e096ad-5d63-4d26-8431-b37d1d41f4d0 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:15.996833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:15.996833Z digest=sha256:c067867865f5cff3b255e269ddcf0178e176d2133f24aabca1ee52d0640a0e6c

Observation 0fbc29ad-f07b-41d6-bf44-229c0d1ae69e · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.000594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.000594Z digest=sha256:c6cc74e0a563f0d7898c3b161feb4f6a8435b9cc1b755151cce50f04616069a7

Observation 337fc6be-176a-4f92-b306-3550a81e47ae · outbound

This paper cites Vila: On pre-training for vi- sual language models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Vila: On pre-training for vi- sual language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.005826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.005826Z digest=sha256:4d5584dafa1f54692cf89754515f4678327fbe0ce6b66b9607bee7b0e3ce8ff8

Observation 09028879-9b50-4a91-9990-f96ac8d269d4 · outbound

This paper cites MM-VID: Advancing Video Understanding with GPT-4V(ision).

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.009332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.009332Z digest=sha256:78ae96dba3df14afaaa872ed587683462cebc8c96370fea3c4932a830887990f

Observation 61998029-218c-4659-9462-e1c4b65bd4bd · outbound

This paper cites Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.016974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.016974Z digest=sha256:a3aab4170873f1a00edf36cafd600a5903d184feee7475040d02d2608c9faac7

Observation 292a7dbe-79ce-45b2-baa7-06880edaf48b · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Improved baselines with visual instruction tuning, 2023

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.022233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.022233Z digest=sha256:9b18e6c55d92cf4e315770f9a10822a5a4d80aa054e2cfce0319ac9b6b092676

Observation 081265e9-c841-44e9-b891-b5c739e4c192 · outbound

This paper cites Visual Instruction Tuning.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Visual Instruction Tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.038887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.038887Z digest=sha256:1e64a3b72bc4feae44dd657338b719777c10535ba85eff83a600c6afa0c7fcfa

Observation 8a912c0c-492a-4a84-9d69-346f68122dbe · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Improved Baselines with Visual Instruction Tuning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.090081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.090081Z digest=sha256:30c9c5602d2a16f7c43abec39ee3e9a194f199b43be382f5bf790daeac53eed3

Observation 5bef9819-8ab3-4451-8d33-f350124ceb12 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.146400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.146400Z digest=sha256:66a576e6c68deed97b22b583d353b7cfa816da60fd422542453e7b9948f3861f

Observation 9f675420-c447-4db8-ae8e-8ca49c964033 · outbound

This paper cites Mmbench: Is your multi- modal model an all-around player? In ECCV, 2024.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Mmbench: Is your multi- modal model an all-around player? In ECCV, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.167905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.167905Z digest=sha256:9b07af03b8f1a2f51aa2f59590b8a32f5a6816701f1d038249f81be772f987de

Observation f5c77256-dca2-4127-bd1b-abe53040b7c9 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.172222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.172222Z digest=sha256:c4b99573532641a0e5b3a852e4503930b7f9d81d3bacf57ee8993805e005c280

Observation dfb87a1a-2332-4769-8d31-8c1dc71dc8c7 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.176720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.176720Z digest=sha256:a974c084d6507aa6e1c92b918639a33ca641154a1f23cd600a8b7a65a3260612

Observation 9b2450fe-3d50-4125-b613-72a606a21916 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.180774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.180774Z digest=sha256:5de16909dd8affa69d78bc6c2bef8324d9876321f29ebe5e9c0c92286ff6156c

Observation 7ebe05e8-b82c-4828-ae81-e366e7491c40 · outbound

This paper cites MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.185082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.185082Z digest=sha256:f0dc46faa57be94e8583bfa5e794e63a5fd06cd14e4fb0d39b0323ce4e5c7f54

Observation e7748671-a59c-48c5-8ab4-38f160ab2934 · outbound

This paper cites Large-scale video panoptic segmentation in the wild: A benchmark.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Large-scale video panoptic segmentation in the wild: A benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.189136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.189136Z digest=sha256:769e4c97b9043cba25ab4e2d2a893633edb4afac797f57807df648637f67f271

Observation 82bceed4-766d-4f3b-b424-12ac6e0820a8 · outbound

This paper cites MOT16: A Benchmark for Multi-Object Tracking.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs MOT16: A Benchmark for Multi-Object Tracking

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.192724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.192724Z digest=sha256:a2e28b22a974b9faf2bc9d46d0978b3eb9de24821f745f929735e564a39456af

Observation bce8364c-c7ea-4525-8f9b-551d991afb5f · outbound

This paper cites GPT-4 Technical Report.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs GPT-4 Technical Report

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.196661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.196661Z digest=sha256:9f11d38bd597393b113924c4ee0d3e174a100146ee4335b2246498310f23807e

Observation 4e445485-bdbf-41f9-809f-4d303609d25e · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs DINOv2: Learning Robust Visual Features without Supervision

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.200512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.200512Z digest=sha256:ce9d6afbfdf7f251404c7a4b1a164d90850b6a00846203dbd0a6c7a24411e072

Observation 44f704c7-6405-4818-9e84-6f40d598c3c7 · outbound

This paper cites VastTrack: Vast Category Visual Object Tracking.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs VastTrack: Vast Category Visual Object Tracking

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.204511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.204511Z digest=sha256:27104609b9dc49900e53b0ab30c095d7de43fa97df3d4d8ccd455ca28725cb40

Observation 18d55902-a0ff-41ff-a7a7-43f10da9bb6e · outbound

This paper cites Occluded video instance segmentation: A benchmark.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Occluded video instance segmentation: A benchmark

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.208465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.208465Z digest=sha256:6327da56df53bbe64338fe1886b8f8a751c6a1b367a1024060d893773cab840d

Observation bb51cf99-ca08-4d5d-85d9-fa50e9f7849b · outbound

This paper cites Chatvtg: Video temporal grounding via chat with video dialogue large language models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Chatvtg: Video temporal grounding via chat with video dialogue large language models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.211980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.211980Z digest=sha256:308babeddaa440e0163b592e49db6a20216e02d415129b3cbfd3ebea040c25b3

Observation 8d706e75-afce-498b-ab60-d015d56c5b6b · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Learn- ing transferable visual models from natural language super- vision

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.216340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.216340Z digest=sha256:a5f98d04dca2d2233d26ca3ba45aad75f4e14ea30b02f71857a3a4701b302113

Observation 5a038d13-666f-4f43-9206-827a491c6539 · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.392617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.220364Z digest=sha256:536c84b07310a193f39070536d52e995e3ce01007e1ada90153e09e0ff26488e

Observation 4dd20f79-1bd9-40c9-86b7-ff622992150e · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Glamm: Pixel grounding large multimodal model

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.382001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.224087Z digest=sha256:e05e6b4b161f3f9a32f1bb255681458cbc06d4559117885ad14473c9f0c4cb83

Observation acaa19dc-91b8-4b4c-8d6f-7a9a736a1174 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs SAM 2: Segment Anything in Images and Videos

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.228898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.228898Z digest=sha256:7880d9465ac4bb04f94855a7af3c21891bf917f17d0b06ec1ddcdce96fe2be6f

Observation f8093b0b-6d46-4454-a930-67eba09d1e5f · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.232974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.232974Z digest=sha256:44715d890ba24273453d9b5aed5590a6725306e8f7d65b6291198401ec9a6430

Observation 15e62f27-e382-42df-a63e-5c2e58abe1f4 · outbound

This paper cites Towards vqa models that can read.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Towards vqa models that can read

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.371336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.236760Z digest=sha256:addaa95750311103a573fdcc93d6a5324c111eccdc5e97e5eea82855e80e71b4

Observation 2e04ca02-d7c2-485d-ba2e-3d3bc77af7c3 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Moviechat: From dense token to sparse memory for long video understanding

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.360500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.240236Z digest=sha256:92201d13732c41f57041fd87457aeca9234da4525bce21f63d72b916d5e45736

Observation d508a717-b65a-4bbb-b85a-e51aa83eb981 · outbound

This paper cites ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language Model.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language Model

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:31:16.760989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.244193Z digest=sha256:35fc1bf0f8d573b57f144038779bef703fbeb671e924b45a2373a60ad6d31bfe

Observation 28d53b87-c524-41fd-847e-9e67d4a81f1f · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.249305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.249305Z digest=sha256:e39649038a00a229928b896110c2c107598de489c86413421da85e28bb08de04

Observation 741d05d3-ed3a-43a4-b5c6-86c6c6eef009 · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities, 2023.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Internlm: A multilingual language model with progressively enhanced capabilities, 2023

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.349620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.253114Z digest=sha256:183658f82cebc763f1668da0b63f15517d55d7ecfde22efae781b8c987d7c3b3

Observation 70bfd6b4-b1e2-45ab-8335-0f85daad112a · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.256385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.256385Z digest=sha256:434caaca67e53cb029db98f37156353db8c43f5871dfa72cd9b14e91cf247bdd

Observation 29a3ebcb-727b-4ba6-a42a-a9f142c12e59 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.337790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.260116Z digest=sha256:6ba1598217db05ee6176db918e9538b12a97b9ce9df51440b216644940182006

Observation d1038f15-88fa-4e47-91c3-82bf4777c9d0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.267385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.267385Z digest=sha256:22c2d036729a1373e584abd30f01507d718106d54060b48a260817a137717406

Observation ff64e833-3947-4420-9219-c3f1ecca0f15 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.270532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.270532Z digest=sha256:a3139fd562412e2c3281292e5c9b5a3879fc10f3f22f3bc804d9339cb0f0a8e2

Observation cddff9e9-2f88-41a1-88f0-c03453586124 · outbound

This paper cites Long-term tracking in the wild: A benchmark.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Long-term tracking in the wild: A benchmark

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.327310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.273637Z digest=sha256:ce8070cd8020f52d6825e15c7c132e934149a3462203b203e26649eac7f93508

Observation d42da109-7df7-40f7-bff1-cb300db98040 · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.276918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.276918Z digest=sha256:f47f5607f9836993241d09e12e6b2960a78d13ec9369afcdbe3130b9bc36265a

Observation 10237edd-c20c-49b6-8230-70466c8f5a45 · outbound

This paper cites Ov-vis: Open-vocabulary video instance seg- mentation.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Ov-vis: Open-vocabulary video instance seg- mentation

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.317369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.280456Z digest=sha256:55490fe409e6fc42061a650da59db051cce29bd1ca786d44ee092a73f9c8923d

Observation 2b3e831b-1d01-41c5-82b8-b4eb10e74db0 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.284252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.284252Z digest=sha256:24e1eb48d039da455bf25fb1fb5c2bca8a1282dd62c4cf7b366011dd1ff612dd

Observation 1767ed95-8310-40fd-8c48-e2186ab2746a · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.288193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.288193Z digest=sha256:6574758012a61b9d87217a0b0bd8230bb6e5c899dc80aba05670251ddab9b7b5

Observation 9f335d46-a268-4bc5-9f39-8e83f6b3e7bc · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.307449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.292409Z digest=sha256:344cc7d600ed264e0d1cc1d704c54dd51c1bd0a05137c853e0bf1204c7cdcc5c

Observation d407407d-f374-4227-96b1-de8c707083f1 · outbound

This paper cites Con- vnext v2: Co-designing and scaling convnets with masked autoencoders.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Con- vnext v2: Co-designing and scaling convnets with masked autoencoders

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.296575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.296402Z digest=sha256:ed13de95244a6f045dbc4085adf6a69d7d63e03e093a80d47b22d8b4a3241f46

Observation 8f71c8b3-5499-45a5-afa6-0eb30d4763b3 · outbound

This paper cites In defense of online models for video in- stance segmentation.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs In defense of online models for video in- stance segmentation

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.286366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.300491Z digest=sha256:92eabd0ed5d081f4ea42f9456b3c5c9c7f52bc27f244622131767563ce195a65

Observation d33eb56b-533a-47a9-ad21-f925e4066063 · outbound

This paper cites ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:31:16.683943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.304478Z digest=sha256:cc05545c4972ffaec811affaf9509c6eec688c16defafc7e469babe5d801d851

Observation 16d49c21-66b4-4731-ba08-7972d6572549 · outbound

This paper cites Pointllm: Empowering large language models to understand point clouds.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Pointllm: Empowering large language models to understand point clouds

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.275803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.310412Z digest=sha256:b03c674b002557446041809d6b6845550ef3483070f488083b9625560c797c82

Observation db653ac8-98b8-4151-a588-ddc956c61c2f · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs xgen-mm (blip-3): A family of open large multimodal models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.316206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.316206Z digest=sha256:b7c2337d65278f062375c03656cb64b3e479ab38085b930c4e4444ee35e7cecf

Observation 92d8533a-d74d-4fc6-9d86-fafc7586d619 · outbound

This paper cites Towards grand unification of object tracking.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Towards grand unification of object tracking

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.264269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.320090Z digest=sha256:8eb34d16302e78e8f4951df94b1a77d9c165155768228b367bcdae9f17a8c14b

Observation 66396e52-a2d5-4f24-a4e7-6a73965e386f · outbound

This paper cites Universal instance perception as object discovery and retrieval.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Universal instance perception as object discovery and retrieval

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.253341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.323879Z digest=sha256:98e2f699f2d1d157024cf1360e911f2efcb55de6744ed7bb9e551f8f5eec017a

Observation aa0dda1e-ef11-48f1-93d8-57b02bde466b · outbound

This paper cites Qwen2 Technical Report.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Qwen2 Technical Report

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.328625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.328625Z digest=sha256:e274fda37a4b6d54bdf9010e162603f88cad9a5900b014f49d41d22b9d211134

Observation e4865195-a39c-45e2-8cd1-a52f834ee2c4 · outbound

This paper cites Video instance segmentation.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Video instance segmentation

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.242680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.332729Z digest=sha256:e149556389acf86a50c6cd1141889f2f4136cfa19d5e057efd2b552465ed5097

Observation fb206a06-e5a0-4e86-87c7-d27abfecece4 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.337149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.337149Z digest=sha256:8b9b3ac457257689a5f077f04c7e05b5d4b466a84939c42532caa68302baf35b

Observation bc93b8b2-4ade-4062-ad9d-44bd2c7f7811 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.341164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.341164Z digest=sha256:801ca8a9929af9af4a430a86021794336e04c6b54cbcbee129ab8f8b8b3bcb30

Observation 8c73889c-c16d-4e12-81cd-b7ad9e9f7d66 · outbound

This paper cites Ctvis: Consistent train- ing for online video instance segmentation.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Ctvis: Consistent train- ing for online video instance segmentation

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.231162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.345081Z digest=sha256:6adb99d91d0a3adcb94f7604a4f9d1cd1038452be231263b56608bf0ca154ba4

Observation 544897a7-4d45-4c5d-8b9d-c1f9c41dced9 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Yi: Open Foundation Models by 01.AI

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.348265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.348265Z digest=sha256:d3fa1f32b4f58f875f7426fe2f86ab3ba600e32571ec86c67cc0008a2fb942a0

Observation e86708ae-6156-472b-a81c-37c15e0a4080 · outbound

This paper cites Bdd100k: A diverse driving dataset for heteroge- neous multitask learning.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Bdd100k: A diverse driving dataset for heteroge- neous multitask learning

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.219853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.351796Z digest=sha256:cfc218c3e94d6f1ec48cfe54afd5c6054d43a60e497940c330bfedf1f4a086ba

Observation 702b7792-0929-453e-b6a4-259a6ce811b7 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.354825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.354825Z digest=sha256:660fef1e32140fd419ccaa7a149d890138042de4e54e5197cf3f3059e9e45005

Observation 6ca3cf87-e249-4e75-85ff-31c580730775 · outbound

This paper cites Osprey: Pixel understanding with visual instruction tuning.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Osprey: Pixel understanding with visual instruction tuning

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:31:17.207662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:31:16.357855Z digest=sha256:e6add7af374a493ed8db4d57c29e667292339200ae062f2c45c3b7c2ead25970

Observation c8d6adaa-7b91-48ae-b276-bb9dd7ce884c · outbound

This paper cites VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.361763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.361763Z digest=sha256:660f5efd089244b6d08c818db355e1dcf6822426672bf7222934c6e05c43dd93

Pith citing papers

Observation b1c18b3a-35d8-48a1-9128-c02701f41d19 · inbound

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models cites this paper.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.658662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.658662Z digest=sha256:dd5ddeda92cfbb6e21ec6928a97c2b876cd685eb65afbc19f217805e8f839e0a

Observation cd2f3850-3373-487a-b4ea-3993548ed273 · inbound

Dense360: Dense Understanding from Omnidirectional Panoramas cites this paper.

Dense360: Dense Understanding from Omnidirectional Panoramas Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:56.252708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:56.252708Z digest=sha256:f7a4a23edf974397d7b274f25b371101b9aa999082e5cbb0790de48beacef040

Observation cab53f23-af9c-4c36-83a5-25d31c6b7cc9 · inbound

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning cites this paper.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:23.293868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:23.293868Z digest=sha256:3b756c6a568d3dec26e7d54bb590f751b178389e185d557e8a04df3674756835

Observation b8a408e2-9f44-4a31-952e-edf973f705e3 · inbound

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding cites this paper.

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.254158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T12:26:01.568507Z digest=sha256:0ba8c894ed372047b3b30c92732130c2c9497f690a2110bbf9ee0d96512a567b