Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2305.11175.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:03:13.068652Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-19T20:28:39.292915Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 62b9628f-a625-40ff-917a-36cafc81723f · inbound
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 43c07343-5c38-44a3-a53b-3a07a7627ca7 · inbound
A Survey on Multimodal Large Language Models VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6365ab3c-6541-497a-8c00-108880ec5a20 · inbound
Kosmos-2: Grounding Multimodal Large Language Models to the World VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5bad6859-bc52-4026-8878-9eb53afb1eca · inbound
A Comprehensive Overview of Large Language Models VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bccf4ac7-ca5c-4fd2-91d4-5bcd947b58f3 · inbound
GPT-Driver: Learning to Drive with GPT VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4d0830bd-85b7-4e67-8aa2-bb964dc523d4 · inbound
Improved Baselines with Visual Instruction Tuning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bb3987c3-7670-4e50-a8b8-b95fa82de8b8 · inbound
MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8dd69875-73c2-416d-9104-4a03f7e2548c · inbound
SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5e7d13db-324a-4214-862d-97aac0223663 · inbound
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 33090c72-7ad7-465e-9114-fb44dccbe226 · inbound
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b4746f15-188b-4783-9fbd-677e6ac04446 · inbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 75d8aeb8-d4d1-4fee-bcc8-209c3b3a9fcc · inbound
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aaae9ad-0aa3-4802-992e-9e3a7360e4e9 · inbound
NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d820e77e-2870-4205-b739-4e60b49df5cc · inbound
MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd6dfdfe-4afe-437f-a1a5-f4490fc688d6 · inbound
DenseMLLM: Standard Multimodal LLMs for Dense Prediction VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c44dd11-9231-4310-9ba9-8b2573aa5f5d · inbound
CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c28589f4-d1d8-47ac-a7a6-d8f8498eefe8 · inbound
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 173
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 293e4431-9be4-413f-97d4-4c252a11b324 · inbound
SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.