Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 53 inbound Pith citation observations for arXiv:2311.07574.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:01:39.976358Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T17:07:25.768424Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation fb08796f-a04b-43cd-99dd-3da7f1d1f638 · inbound
A Survey on Multimodal Large Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 246e4615-ae3c-4a41-b907-cb9847acbc77 · inbound
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dc43a9a1-d6cd-4c80-bb6b-37ce19e8e0e5 · inbound
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3ced92d1-13d4-4817-80a1-a4ac61a6905e · inbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 846e3349-49b1-4036-be1c-66763a84d450 · inbound
DeepSeek-VL: Towards Real-World Vision-Language Understanding To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 41079ee0-4245-4c6d-95d0-d2ae23429ab5 · inbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 91f73498-8512-43d5-b6a2-cc4d57037c79 · inbound
Are We on the Right Way for Evaluating Large Vision-Language Models? To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e4c6ebed-d918-48a8-a469-81cab83b8464 · inbound
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bdff8938-636e-4351-b0f2-108942cf6752 · inbound
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9be80528-7c79-4df7-a9ac-42d7cf12cd4e · inbound
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ec5e6333-e0bf-4e44-8a52-61b818a22b4c · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3b444aa8-0c49-43fd-906d-460b82451ffc · inbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cba2aec2-877a-45a6-8f82-dbfb08cdcf5b · inbound
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37c6eede-1eec-4593-8dac-5d4cdcd87a11 · inbound
BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 124
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2353e9f8-d74d-4bf5-8e8e-4b59da94f0f5 · inbound
Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fb38b55-2166-4303-89fe-b62a5283848a · inbound
Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f63519f4-1867-4e76-b350-9d23eca62241 · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 245
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c14f66df-556f-4525-b716-5a3e9cc4b655 · inbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbd852c3-0825-4781-914e-37e6b6a16722 · inbound
V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 120
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbefc10c-0c29-47b3-a559-8f547d0cbcfa · inbound
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 35b58c34-35a2-4dd1-8441-376a04fccdd3 · inbound
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4151c0ed-3e43-4811-8928-2af738694700 · inbound
Visual Large Language Models for Generalized and Specialized Applications To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 258
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a81bae1f-746a-4107-90b0-ebf99c150d8f · inbound
Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70518c8b-093a-44e4-9bac-08f2e0062082 · inbound
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 126
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5d65e6f-ce52-4940-894f-d073327309c7 · inbound
MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6280041-f4b1-4ef6-ae4a-877331f0c2b9 · inbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d87f0295-6e5f-4975-a7dc-91dd263cd06e · inbound
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8b9c1474-9b5a-4f77-9bfe-1f1c04b2e7e5 · inbound
Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04a819bd-1fc3-4b4d-8bc9-74b786916413 · inbound
Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9fbeeaa-147d-48ab-8af4-d1dc59b836ec · inbound
Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 380b9a90-ab99-4a3d-8c4d-52846bee1ac0 · inbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24283774-8d39-46e8-b0b9-60d9db83d88c · inbound
FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 136
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00b24c17-26cc-4e54-9912-f809d9449f77 · inbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1df6994e-e983-4736-9c96-d1dc5913f244 · inbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa307276-50eb-488a-b3c3-61cad5c90d37 · inbound
EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 773e5efe-48b2-4f08-824e-8193d0d72480 · inbound
Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 464049c1-2cff-40ee-a310-a8cab0779050 · inbound
SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d658a2d-493f-4c3d-8b4e-c7ef1a5e2a14 · inbound
Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e37eb4b-4166-45b0-b6df-e75df5fd21a9 · inbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b970ecbc-9311-4a1b-bd3c-f83feadd74b4 · inbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95153f93-0b08-4be7-a40f-5c9f3ef25d0c · inbound
See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28cd8b76-9dd8-4a98-94ff-bdf1db514fcc · inbound
Metadata Management for AI-Augmented Data Workflows To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8d5f0d5-a4b0-4106-a68a-5824764c9cde · inbound
UItron: Foundational GUI Agent with Advanced Perception and Planning To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97f328e9-a648-4b06-ba86-243fe4c8799c · inbound
Improving Large Vision and Language Models by Learning from a Panel of Peers To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 260f681b-8edd-4671-8ead-82b65f1cc844 · inbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ce4d7eb-816a-44b9-acb1-7698013bf443 · inbound
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0174631a-5b6f-4a47-8001-62f592d94ac9 · inbound
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 264
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2de9cc4a-c6b4-43ab-92a2-701dbd88eb47 · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 295
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3bc2e407-a963-4718-b5a0-7862c86e6e55 · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 295
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 31f9f1ea-8638-4cf0-88e6-79b6c128692d · inbound
Infinity-Parser2 Technical Report To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 011eca1f-ffd9-45a6-a8a8-4b041fdf3c3c · inbound
Infinity-Parser2 Technical Report To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb964605-00b7-4897-b627-d49e8787f565 · inbound
Twins: Learn to Predict Unified Representations with Focal Loss To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 281
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39244b89-6a3e-4823-8faa-1fa97e178d37 · inbound
DistMoE: Private-data Rehearsal-free Routing in Mixture-of-Experts for Distributed Instruction Tuning To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.