Pith. sign in

Paper Citation Record · LEDGER

VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2305.11175.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.11175 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:35:52.101809Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T20:28:39.292915Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 62b9628f-a625-40ff-917a-36cafc81723f · inbound

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models cites this paper.

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:34.186308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T20:25:33.854923Z digest=sha256:73f8df349a9b276b59b53c408ff4af8add4526252c643de90ceb508fa8fcfbd1

Observation 43c07343-5c38-44a3-a53b-3a07a7627ca7 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.616141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:9bb317eaa30177e36bf5097561a8629be662b5ceadd0018f66aae8686c4307e2

Observation 6365ab3c-6541-497a-8c00-108880ec5a20 · inbound

Kosmos-2: Grounding Multimodal Large Language Models to the World cites this paper.

Kosmos-2: Grounding Multimodal Large Language Models to the World VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:19:48.098848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T05:19:47.907355Z digest=sha256:28f23ccf84a65c30ab955c81f6f5eea2e9f66b96363b4352d48fcd32fb3050cb

Observation 5bad6859-bc52-4026-8878-9eb53afb1eca · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.294606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:52c29ae815869cce1fb50f6e32ad111b8a5e72bdbbdc37a47881a4f6a69dfac5

Observation bccf4ac7-ca5c-4fd2-91d4-5bcd947b58f3 · inbound

GPT-Driver: Learning to Drive with GPT cites this paper.

GPT-Driver: Learning to Drive with GPT VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:05:32.007128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T15:05:31.928650Z digest=sha256:02285c8ae107d20965d1234db198c10ba1565b36d88f7a5d3dfbddd76aa5a471

Observation 4d0830bd-85b7-4e67-8aa2-bb964dc523d4 · inbound

Improved Baselines with Visual Instruction Tuning cites this paper.

Improved Baselines with Visual Instruction Tuning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:11:33.951814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T19:11:33.783746Z digest=sha256:7048c916dc869e3c25b354893a6448a7cb767b4a5d6abf72056abe01e0b8a32b

Observation bb3987c3-7670-4e50-a8b8-b95fa82de8b8 · inbound

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning cites this paper.

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:13:08.994163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T07:13:08.867745Z digest=sha256:3aab589a9533b24859c8b2eca38a9cc1fad8e687d200bbc9d552e8821ed202a9

Observation 8dd69875-73c2-416d-9104-4a03f7e2548c · inbound

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models cites this paper.

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:03:27.007139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T03:03:26.723464Z digest=sha256:f1530ca3ecc47e13cda83e01eb33d793eb9a8ab8de4f11b3c49c57eb4c78f09d

Observation 5e7d13db-324a-4214-862d-97aac0223663 · inbound

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents cites this paper.

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:09:46.610091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-17T10:09:46.447508Z digest=sha256:a5958c50d4fc67b3bddc4791f960afdddfbc798e555f3b1c8e7b684326c7e309

Observation 33090c72-7ad7-465e-9114-fb44dccbe226 · inbound

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models cites this paper.

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:33:30.420855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T02:33:30.143907Z digest=sha256:b18338e3b9ea28d1155a17f2b99cbaaf8ee56f9638a28bdfdc31a1651f451700

Observation b4746f15-188b-4783-9fbd-677e6ac04446 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 115

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.310691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:3796f1b26ce0598dc77c2459c0a931bd233101c33cf5cbbce19b4535786eaa61

Observation 7ea145d0-83e8-4646-9924-11a33d3549e1 · inbound

The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use cites this paper.

The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T19:47:48.239370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:47:48.239370Z digest=sha256:da965a8078db235a5857d8051d394a399a6bbd4048e90e9895d4d372f662fab8

Observation 572d61f9-5bd5-42ae-aa6d-732604ef6548 · inbound

CCExpert: Advancing MLLM Capability in Remote Sensing Change Captioning with Difference-Aware Integration and a Foundational Dataset cites this paper.

CCExpert: Advancing MLLM Capability in Remote Sensing Change Captioning with Difference-Aware Integration and a Foundational Dataset VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T18:41:08.137777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:41:08.137777Z digest=sha256:5d2a9ceffa002ab8e64cf13ab68470ac6fa3df92469eb3e2f56bf18e6a67bfe4

Observation 9096e147-9792-4ea6-8cbb-84a8b06d9cec · inbound

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model cites this paper.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.049912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.049912Z digest=sha256:5f0495ce5ee1ac4bfc6af50b7fc848d56ce7e7506613e9ebaaabed44de09c045

Observation 4e420783-569b-4526-88f5-4d07aea7ae3d · inbound

AutoLife: Automatic Life Journaling with Smartphones and LLMs cites this paper.

AutoLife: Automatic Life Journaling with Smartphones and LLMs VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T11:12:32.992296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:12:32.992296Z digest=sha256:02199e71b656cf8faacee5daeb62228aef874ae0843ebed73418dd76b622a77c

Observation 75d8aeb8-d4d1-4fee-bcc8-209c3b3a9fcc · inbound

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis cites this paper.

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T20:03:13.068652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:03:13.068652Z digest=sha256:16586a02c729662068530f5a01ff204246b0170a9e1860a5d5df10760e8b609e

Observation 6aaae9ad-0aa3-4802-992e-9e3a7360e4e9 · inbound

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning cites this paper.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.348314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.348314Z digest=sha256:eae668bc44178a9157964835a93ba92b2dae8493ca421d8db6f8cd0571839bc6

Observation d820e77e-2870-4205-b739-4e60b49df5cc · inbound

MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation cites this paper.

MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T14:50:57.945485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:50:57.945485Z digest=sha256:81ffe52e1fe903e795b8809f053964101f40134001c50cdd2d7ddb7977e38c26

Observation bd6dfdfe-4afe-437f-a1a5-f4490fc688d6 · inbound

DenseMLLM: Standard Multimodal LLMs for Dense Prediction cites this paper.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.980496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.980496Z digest=sha256:7a47d0667395ee3ff807c83fea7355193f0eced5372b7d06d664098f8c79bb7f

Observation 6c44dd11-9231-4310-9ba9-8b2573aa5f5d · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:17.122294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:c5268a2a98371dbab5ec6475ca3fd0e8778b972418ca0e2339f1f74e28b6ca59

Observation c28589f4-d1d8-47ac-a7a6-d8f8498eefe8 · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 173

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:09.601368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:9c47314faf08f1a8fc18d396d9827e731c173d5f82a70b94d79a91d5072b57a9

Observation 293e4431-9be4-413f-97d4-4c252a11b324 · inbound

SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning cites this paper.

SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:10:26.578809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T13:08:36.004520Z digest=sha256:6f816365e407401c3e30bd63b22aa5dc9246641180f461ac51a8dba89b643050

Observation af2ea5ac-ae00-467d-a394-db444c61cbf8 · inbound

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models cites this paper.

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 210

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:52.101809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:35:52.101809Z digest=sha256:6ae629da1b9d67aff37e66ff93207f52cdd0024be01c8791d5028b757ca14995