Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2306.17107.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T17:13:26.039380Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T15:39:56.491911Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation d4bfc8e1-85d5-47e8-bb2f-fba9bc7fdeb5 · inbound
Otter: A Multi-Modal Model with In-Context Instruction Tuning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f0d2089-6468-4dae-8b88-6b1cde9056d4 · inbound
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f81cb006-3de7-4356-a7d2-478facbfd48b · inbound
Improved Baselines with Visual Instruction Tuning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3d5da0f1-f352-4fc5-ad69-54e92950cc89 · inbound
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 752f4153-e37b-4e86-9fd2-0590c19295c6 · inbound
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 183
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0ee7bdb2-737a-4bac-bedd-c87c5c5dc7ec · inbound
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation da761b21-bc1c-4309-9a6b-0202105cea7a · inbound
Yi: Open Foundation Models by 01.AI LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 40490174-7897-4290-ba07-fe31d5bcba88 · inbound
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7961a538-9c52-4f4e-b58f-0293a8b8c28b · inbound
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 149
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a0c9ca41-5014-4f37-9a4e-26ea794e962c · inbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 43929916-c215-4956-965e-51a3ddd062bb · inbound
LLaVA-OneVision: Easy Visual Task Transfer LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7d29eb95-245f-4fdb-a75d-d9ca6d406c35 · inbound
NVILA: Efficient Frontier Visual Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0ade628d-a4e4-4dfa-9dd1-ceeb3e3f2dc9 · inbound
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a793304d-b92d-4fa4-8199-35015e9dc1ae · inbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e83eed00-f88e-40ce-8d5a-061b7c511f41 · inbound
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b4f09d17-eaf0-4993-bfdf-572bab3e51df · inbound
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a346cf23-6cad-4a8b-867c-efcd33e76a44 · inbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6e354e6-b377-4bf1-90e4-2954bd1181c7 · inbound
FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89728534-4ffa-413d-875c-af011ef21c29 · inbound
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5a0e74d-ac7d-4915-8c45-3e16b5a9adaa · inbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 109
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6a9aadf-9721-470a-a358-3ea30994739c · inbound
CoMemo: LVLMs Need Image Context with Image Memory LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 113
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccec615d-d13e-4de7-ba4d-53c35f5b8de5 · inbound
Synthetic Visual Genome LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b24f0e5f-e49b-4afd-83d0-bc32bddfb148 · inbound
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 122
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a96f1011-b795-4124-9646-29a6fe338376 · inbound
Multimodal Mathematical Reasoning with Diverse Solving Perspective LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 568769e4-9114-4bd8-9e2e-fe8d9b7c8914 · inbound
Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40b253b0-067d-4f80-b54f-90b17167e97a · inbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a888602-e291-42c8-9625-1461a92fcc8a · inbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b8f319d-4258-4591-95f7-2899fd3f623a · inbound
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e5af482a-0791-40c7-92f8-9b122181d297 · inbound
CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f7999db-19ce-4621-af6c-b8935cf10a27 · inbound
Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bbfdfc6-adfc-4968-8047-9ba9f6bab163 · inbound
OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 000e01be-5b7f-4368-870a-4cb39fd6e76e · inbound
Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d7ab97c0-7887-4c80-bb99-5e36e3e042cc · inbound
Closed-Form Spectral Regularization for Multi-Task Model Merging LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b85f4393-3752-4c7b-afd8-3bd4d3ec02eb · inbound
Vision Language Model Helps Private Information De-Identification in Vision Data LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c5bd9300-5a26-42df-9cc6-136b81cfbea2 · inbound
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 228
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 422eb2e9-c7f5-49af-a622-94650ef89f12 · inbound
GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 929cea13-7215-4400-b797-b5db081d1334 · inbound
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 135
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7a7a505-b829-4e90-9f55-e476e8698b32 · inbound
Twins: Learn to Predict Unified Representations with Focal Loss LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 282
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af204b94-4cd4-4e3f-91de-fc9cdbedc8e3 · inbound
ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cab4f00-62f7-40d1-9b1f-c814d20e9488 · inbound
PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.