Pith. sign in

Paper Citation Record · LEDGER

Are MLMs Trapped in the Visual Room?

As of 9 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2505.23272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23272 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:54:27.130704Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T08:20:56.561166Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T08:22:10.936561Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 10197d36-6930-412d-943b-607fa623afb5 · outbound

This paper cites Advances in neural information processing systems35, 23716– 23736 (2022).

Are MLMs Trapped in the Visual Room? Advances in neural information processing systems35, 23716– 23736 (2022)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:28.081107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:54:24.966379Z digest=sha256:5042d709e33dfac2011d69546a203159c050327d79cc7ebc285b946574c10f5e

Observation 6a90386f-f6b1-4cf6-9e9f-6d65c061194c · outbound

This paper cites Qwen2.5-VL Technical Report.

Are MLMs Trapped in the Visual Room? Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:25.054856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:25.054856Z digest=sha256:ba30cc789d74c348dee092a129011e779ed344593edb54566273fed731d29043

Observation 3b4ddf6d-1c1e-4d6f-9e42-8b6807b3a149 · outbound

This paper cites Towards Multimodal Sarcasm Detection (An _Obviously_ Perfect Paper).

Are MLMs Trapped in the Visual Room? Towards Multimodal Sarcasm Detection (An _Obviously_ Perfect Paper)

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:25.159760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:25.159760Z digest=sha256:fbead7ca272a0332f7aa53e10799a18f80b17f0770b603946b1c711a866a3b1d

Observation 5139d084-0e47-43d3-a0c9-ec38f6bb92da · outbound

This paper cites Advances in Neural Information Processing Systems37, 110805–110853 (2024).

Are MLMs Trapped in the Visual Room? Advances in Neural Information Processing Systems37, 110805–110853 (2024)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:27.926495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:54:25.251838Z digest=sha256:6cb19b0ce306a607a0396abe6efd376253cd9e5d1ed46d5d4ff4c0869a6e6195

Observation 19863e85-f32f-4496-a8aa-4b9b282be700 · outbound

This paper cites an unresolved cited work.

Are MLMs Trapped in the Visual Room? Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:25.350577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:25.350577Z digest=sha256:c93f35323e9ca83b4ddbfe847bc61cbbb044c4f1bef2f7899fd1363554a011ce

Observation 59ba3b7d-bae5-4f2d-9f27-a7b0bdcedd10 · outbound

This paper cites GPT-4o System Card.

Are MLMs Trapped in the Visual Room? GPT-4o System Card

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:25.466576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:25.466576Z digest=sha256:2bc313f2d3509cca1971879727e3677c3d2bf75f148ebb6b5962df3e451ac040

Observation 02bc6d2b-074d-4321-8b2e-9ef466e960d9 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Are MLMs Trapped in the Visual Room? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:25.598847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:25.598847Z digest=sha256:8c73aa00884c1a5877e89760f781ed16fcb3e8d20bf66f4cb13ebcf25a30794c

Observation fb71c79b-557a-4685-8682-775a215e124a · outbound

This paper cites In: International conference on machine learning.

Are MLMs Trapped in the Visual Room? In: International conference on machine learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:25.728615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:25.728615Z digest=sha256:dd3e67eae72085dad8a79c2f9f555f8ab48f699104301eb196316a9299559527

Observation 9eb782d3-f035-447f-843a-565f68f4b0ac · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Are MLMs Trapped in the Visual Room? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:25.856176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:25.856176Z digest=sha256:bb83459284b16f4cf784b551abc19b40cd34f50b45ce1fa579c71e3dbcc1d3f8

Observation e29d5e5b-71d6-414e-9023-a1cc80353618 · outbound

This paper cites Generated Knowledge Prompting for Commonsense Reasoning.

Are MLMs Trapped in the Visual Room? Generated Knowledge Prompting for Commonsense Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:25.942571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:25.942571Z digest=sha256:23277629a5eb916992c56ceb281b17e69a50862dbcb9fa950eb0f38b2a5792b2

Observation 0d3d752f-9ee3-401d-9e15-2d0d2aa687e2 · outbound

This paper cites In: Proceedings of the 32nd ACM International Conference on Multimedia.

Are MLMs Trapped in the Visual Room? In: Proceedings of the 32nd ACM International Conference on Multimedia

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:27.742685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:54:26.039290Z digest=sha256:c97d2e710122e20d27956d1397115f646baa5d0f977d9d055f05ad3bdc30f454

Observation d7c945a3-792b-4025-a7e2-1e1504c5284a · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Are MLMs Trapped in the Visual Room? DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:26.194743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:26.194743Z digest=sha256:b6258e5b48e1bc48b3a4a29a74d6c9739f2146d246f23e91d752d3e4d4fcb86d

Observation 68e865bb-5934-4a6b-ad97-0c9996bb4c98 · outbound

This paper cites MMSD2.0: Towards a Reliable Multi-modal Sarcasm Detection System.

Are MLMs Trapped in the Visual Room? MMSD2.0: Towards a Reliable Multi-modal Sarcasm Detection System

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:26.308818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:26.308818Z digest=sha256:70a33f108630ee0f81a7ea3c8021e2e00b13c6cbc8e0daae6a3008209088a6da

Observation baee13cc-5c43-43cc-b383-46088c667f8e · outbound

This paper cites In: International conference on machine learning.

Are MLMs Trapped in the Visual Room? In: International conference on machine learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:26.367430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:26.367430Z digest=sha256:7b0362c2e33107ef207c5e91f20a0eadfa7cff86b0e1879c86f0b9cd5696377a

Observation 7b9972cb-ffeb-4639-9573-b00babf8f469 · outbound

This paper cites Scholarpedia4(8), 3100 (2009).

Are MLMs Trapped in the Visual Room? Scholarpedia4(8), 3100 (2009)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:27.597716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:54:26.439968Z digest=sha256:8ff5dc51e291ecf71726e036dadb8fc4e5ac6d85fb993e1c663c28ce4193d08e

Observation f815efce-317d-4461-a7f1-7c8f898e212d · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Are MLMs Trapped in the Visual Room? LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:26.520235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:26.520235Z digest=sha256:14e4e85ade1f806381f71510a9bdba3ea175defb49f9d49ea9aa0aaa70d04727

Observation 7e8d331a-8437-4602-8717-6fa5d7c5c7e3 · outbound

This paper cites Information Fusion 103, 102132 (2024).

Are MLMs Trapped in the Visual Room? Information Fusion 103, 102132 (2024)

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:26.649855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:26.649855Z digest=sha256:243e7c059d37e0e5ec5069a553dfc9c7713d3e006a0824acf0726d612fbe26e3

Observation 00371fb9-2991-4069-9f45-b31c4e4dd40c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Are MLMs Trapped in the Visual Room? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:26.702049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:26.702049Z digest=sha256:0eeba599516ddf890b979986174484ba042873f933a847f1cd9c3df0702557fa

Observation e8df1fdf-8a55-4fe2-8821-1465ea6aeebf · outbound

This paper cites Can GPT-4V(ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis.

Are MLMs Trapped in the Visual Room? Can GPT-4V(ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:26.786207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:26.786207Z digest=sha256:929b22ff68d39760fc97f3da872719214b91962a12bb61c58c17dcbba04c1f5f

Observation 7456e1fd-f150-46b0-9bba-6f41c97a0ce7 · outbound

This paper cites DialogueLLM: Context and Emotion Knowledge-Tuned Large Language Models for Emotion Recognition in Conversations.

Are MLMs Trapped in the Visual Room? DialogueLLM: Context and Emotion Knowledge-Tuned Large Language Models for Emotion Recognition in Conversations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:26.879990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:26.879990Z digest=sha256:a7a79286e91ba33387c2f90258753c436bcde0dc90bfb967174f54f3214e3938

Observation d057ae2d-c8ef-41de-b82c-769ac8f6f1c4 · outbound

This paper cites Advances in Neural Information Processing Systems36, 18794–18805 (2023).

Are MLMs Trapped in the Visual Room? Advances in Neural Information Processing Systems36, 18794–18805 (2023)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:27.444845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:54:26.971368Z digest=sha256:c00037cbbc013ab42b99c594539d9372a4f3a718aa7a63a20dc37377cbfe778f

Observation 9279b93b-17e7-469f-a099-4376b71c4232 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Are MLMs Trapped in the Visual Room? Multimodal Chain-of-Thought Reasoning in Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:27.034519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:27.034519Z digest=sha256:9aefbe80e78ee112c8f3d3e9be78251da0b687177656d0bcb2b6b0941fcc4b0f

Observation 38fa8348-8c66-410d-89c0-fb7767f6dca1 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Are MLMs Trapped in the Visual Room? MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:27.130704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:27.130704Z digest=sha256:6b5471698f0be02172d1b3360018e3a724668a08f7c7615b0811e3be31094df6

Pith citing papers

Observation 51b99c23-2f2c-480c-9eaf-133d03052004 · inbound

Commander-GPT: Dividing and Routing for Multimodal Sarcasm Detection cites this paper.

Commander-GPT: Dividing and Routing for Multimodal Sarcasm Detection Are MLMs Trapped in the Visual Room?

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:22:10.940154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T08:20:56.561166Z digest=sha256:b45dec273725a2b76c3e69ffd60447a0fd377971e49b75356e97694f316a79e5