Pith. sign in

Paper Citation Record · LEDGER

VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2305.11175.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.11175 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:03:13.068652Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T20:28:39.292915Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 62b9628f-a625-40ff-917a-36cafc81723f · inbound

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models cites this paper.

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:34.186308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T20:25:33.854923Z digest=sha256:bed9cf5231383fbeab617fb2e35eb14191dd11ea6695a2e38dd43d2ab0114ecc

Observation 43c07343-5c38-44a3-a53b-3a07a7627ca7 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.616141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:d83facce6d8fe3ddb97a97312e83b8187f7c65e7f34f988d3952c9f78e39017d

Observation 6365ab3c-6541-497a-8c00-108880ec5a20 · inbound

Kosmos-2: Grounding Multimodal Large Language Models to the World cites this paper.

Kosmos-2: Grounding Multimodal Large Language Models to the World VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:19:48.098848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:19:47.907355Z digest=sha256:94edf2564a818f11a9c294fc5e7eb6df52eef0af5957091614e19dd21493222a

Observation 5bad6859-bc52-4026-8878-9eb53afb1eca · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.294606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:005a2b58a99537f0dabd93091ba93f444164ba39279b82490d73bfea450d268b

Observation bccf4ac7-ca5c-4fd2-91d4-5bcd947b58f3 · inbound

GPT-Driver: Learning to Drive with GPT cites this paper.

GPT-Driver: Learning to Drive with GPT VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:05:32.007128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T15:05:31.928650Z digest=sha256:5b4781a8d7eddfadd62147da2ee7ccd31e4d36e6f7184fba8278e54d812501c6

Observation 4d0830bd-85b7-4e67-8aa2-bb964dc523d4 · inbound

Improved Baselines with Visual Instruction Tuning cites this paper.

Improved Baselines with Visual Instruction Tuning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:11:33.951814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T19:11:33.783746Z digest=sha256:77a7a0a32ad0c2f50d21872fab7c370e63c88fe1a89cc6d538a598f94ab4dd77

Observation bb3987c3-7670-4e50-a8b8-b95fa82de8b8 · inbound

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning cites this paper.

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:13:08.994163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T07:13:08.867745Z digest=sha256:1e7f1da18ae13bdda799a700ecd1cb2ad86b0eac6f25cb7a78b08346d9838d5f

Observation 8dd69875-73c2-416d-9104-4a03f7e2548c · inbound

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models cites this paper.

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:03:27.007139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T03:03:26.723464Z digest=sha256:2749b9255b2e5d487735a94498d004e6d9c7614054f8de3d062a32c4c1e64b15

Observation 5e7d13db-324a-4214-862d-97aac0223663 · inbound

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents cites this paper.

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:09:46.610091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-17T10:09:46.447508Z digest=sha256:093e9545dec2ac517b4765d9145979e7e37a6d1773171ed97d2f271428875044

Observation 33090c72-7ad7-465e-9114-fb44dccbe226 · inbound

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models cites this paper.

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:33:30.420855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T02:33:30.143907Z digest=sha256:7c646a5f01d4a182e2ae7d1e041f72bf9b4b34a2b5bfbd2c644ae9b44cc8b761

Observation b4746f15-188b-4783-9fbd-677e6ac04446 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 115

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.310691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:b6c3e259c190a6c41f45e4cedd82550038048bcb120cf0e8d703f75a4b5347bb

Observation 75d8aeb8-d4d1-4fee-bcc8-209c3b3a9fcc · inbound

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis cites this paper.

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T20:03:13.068652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:03:13.068652Z digest=sha256:894ae05871b98188f34f82a0d05119df1f8d8e68fa30a30afe5b999152339afe

Observation 6aaae9ad-0aa3-4802-992e-9e3a7360e4e9 · inbound

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning cites this paper.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.348314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.348314Z digest=sha256:2beaf26ae237c284d0444eab919f1d9017c7341158910ff4301d092454eed87e

Observation d820e77e-2870-4205-b739-4e60b49df5cc · inbound

MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation cites this paper.

MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T14:50:57.945485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:50:57.945485Z digest=sha256:806942e5728d3bf467e7b358ff90913bb08a891c4aeec5d79a826e545f6cadb6

Observation bd6dfdfe-4afe-437f-a1a5-f4490fc688d6 · inbound

DenseMLLM: Standard Multimodal LLMs for Dense Prediction cites this paper.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.980496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.980496Z digest=sha256:264d814ed2bd7b0cb568e02667e3b3bfdeac20256a579acaf1d0411b0750ad70

Observation 6c44dd11-9231-4310-9ba9-8b2573aa5f5d · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:17.122294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:0c8e1b2e5933a48ec85d26d10fe6b96dd18820b887f8996785a519b17ab79720

Observation c28589f4-d1d8-47ac-a7a6-d8f8498eefe8 · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 173

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:09.601368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:0b0c0e212eba1be4bf2ae89e12e866c493287da96264a091aefc0f50bd5ac0a2

Observation 293e4431-9be4-413f-97d4-4c252a11b324 · inbound

SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning cites this paper.

SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:10:26.578809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:08:36.004520Z digest=sha256:2f82f4fe6e49b2c0d7195a9f63c26c77f893f368bcd3ea16755f5e715ed86bf2