Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:41:07.654957Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2412.10758.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:41:07.654957Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation de7930ef-f9ba-4cdf-8795-e67183fa8da3 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Unveiling Encoder-Free Vision-Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f3c027a-23f5-49e5-bf7b-c9340a66c595 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29fe954b-93b0-43eb-ba0b-5604e6484cf9 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Visual in-context le arning for large vision-language models,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c6a1830-2b1b-448d-9a27-2f52f892de80 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Claret: P re-training a correlation-aware context-to-event transformer for eve nt-centric gener- ation and classification,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 884311d6-0dfb-43ad-acea-83f94f7efda6 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Eventber t: A pre- trained model for event correlation reasoning,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f360099-d185-4b02-80c6-1149ae41d554 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Visionllm: Large language mode l is also an open-ended decoder for vision-centric tasks,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c35f26c-cc7b-4475-89e5-dd16497b03d6 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2fb4fb8-c10f-4b83-9891-1ae0033fe3ed · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dffc29d3-fe2f-412e-94bb-ab98399cc140 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Enhancing Large Vision Language Models with Self-Training on Image Comprehension
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efdb86d5-6a57-4f20-8773-6fa05cfdb689 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69dd160d-b850-46cc-9e8e-b2abe60ea223 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Triple sequence generati ve adversarial nets for unsupervised image captioning,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79da6fd4-c041-4c53-b86b-d7ba60b2fd29 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Sketch storytelling,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08ee393c-8162-45b1-a029-1bcd41085a7a · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d269bcc-721f-4cba-ab9b-9803bb332f29 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47267141-3a43-4ce5-9677-5de5bbb57551 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e08b26e-817c-4f41-9012-f37a34ff83ae · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Multimodal event transformer for i mage-guided story ending generation,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 290f8d76-452e-4d2d-bfc2-4da99cfc6534 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa2f60fb-8b6b-4693-8d29-832a263aaa79 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Vilt: Vision-and-language t ransformer without convolution or region supervision,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5b40c818-fa1b-4aaa-a51c-8228ed812c57 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Style-aware contrastive learning for multi-styl e image caption- ing,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d9efcc2f-8978-4d23-97a8-97fad21a7f40 · outbound
Optimizing Vision-Language Interactions Through Decoder-Only Models Unveiling Encoder-Free Vision-Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.