Pith. sign in

Paper Citation Record · LEDGER

Optimizing Vision-Language Interactions Through Decoder-Only Models

As of 12 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2412.10758.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10758 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:41:07.654957Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de7930ef-f9ba-4cdf-8795-e67183fa8da3 · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

Optimizing Vision-Language Interactions Through Decoder-Only Models Unveiling Encoder-Free Vision-Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.405342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.405342Z digest=sha256:cdbec4e439e9f0b4c4e5419eafbf1c57cbf7b1a44e4baf7394d89ec7dbf54a52

Observation 0f3c027a-23f5-49e5-bf7b-c9340a66c595 · outbound

This paper cites Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models.

Optimizing Vision-Language Interactions Through Decoder-Only Models Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.418467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.418467Z digest=sha256:0acfc58107a09eb741ba61db90a79331cd47b93328f99888a8d199203f7d6ec5

Observation 29fe954b-93b0-43eb-ba0b-5604e6484cf9 · outbound

This paper cites Visual in-context le arning for large vision-language models,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Visual in-context le arning for large vision-language models,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.424605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.424605Z digest=sha256:115839a57211db0a4affe4fb59d25a3d15f26d40402bc4512bb648065ea2d156

Observation 7c6a1830-2b1b-448d-9a27-2f52f892de80 · outbound

This paper cites Claret: P re-training a correlation-aware context-to-event transformer for eve nt-centric gener- ation and classification,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Claret: P re-training a correlation-aware context-to-event transformer for eve nt-centric gener- ation and classification,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:08.041191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:07.434018Z digest=sha256:942a8b57fbc0680565db53f30fa0393250c2983b55807a0917af3080be358a88

Observation 884311d6-0dfb-43ad-acea-83f94f7efda6 · outbound

This paper cites Eventber t: A pre- trained model for event correlation reasoning,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Eventber t: A pre- trained model for event correlation reasoning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.439801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.439801Z digest=sha256:d7e20c1e71b432ea0c85895ebdb92273171bd23cacfd22ce1609b25d4fea76d8

Observation 0f360099-d185-4b02-80c6-1149ae41d554 · outbound

This paper cites Visionllm: Large language mode l is also an open-ended decoder for vision-centric tasks,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Visionllm: Large language mode l is also an open-ended decoder for vision-centric tasks,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:08.006708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:07.445111Z digest=sha256:eb0d874321096826e536c9930b6d10308ac2417c2c7e22e05ca9833eb2feb7ca

Observation 8c35f26c-cc7b-4475-89e5-dd16497b03d6 · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

Optimizing Vision-Language Interactions Through Decoder-Only Models VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.450255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.450255Z digest=sha256:43082fcd084e30a188fb79c0ca15cfc86742623220403fe74e84708192f9b069

Observation d2fb4fb8-c10f-4b83-9891-1ae0033fe3ed · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Optimizing Vision-Language Interactions Through Decoder-Only Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.456101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.456101Z digest=sha256:e410bddd1ae62f01e6b374b67e593c0a173555357d7604de42be010686ee8aa8

Observation dffc29d3-fe2f-412e-94bb-ab98399cc140 · outbound

This paper cites Enhancing Large Vision Language Models with Self-Training on Image Comprehension.

Optimizing Vision-Language Interactions Through Decoder-Only Models Enhancing Large Vision Language Models with Self-Training on Image Comprehension

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.461270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.461270Z digest=sha256:27ec8dadc101dd36cfb2573a89d39303457f75d90564868d2044884c37104557

Observation efdb86d5-6a57-4f20-8773-6fa05cfdb689 · outbound

This paper cites Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning.

Optimizing Vision-Language Interactions Through Decoder-Only Models Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.467266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.467266Z digest=sha256:3b46160e4e7a6e94e5cd2e82fe7bdf5a38ace8b062614f8193a29d0b7db5e277

Observation 69dd160d-b850-46cc-9e8e-b2abe60ea223 · outbound

This paper cites Triple sequence generati ve adversarial nets for unsupervised image captioning,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Triple sequence generati ve adversarial nets for unsupervised image captioning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.473518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.473518Z digest=sha256:d9653256e526ce61f3a3758da17d5aa0f9e76d7273e9113dcd4a1ccecab2f72e

Observation 79da6fd4-c041-4c53-b86b-d7ba60b2fd29 · outbound

This paper cites Sketch storytelling,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Sketch storytelling,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.605049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.605049Z digest=sha256:50db326e9a7a0a47e2670f20883c1192124a7c2054d053f084ec64d3b8554789

Observation 08ee393c-8162-45b1-a029-1bcd41085a7a · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Optimizing Vision-Language Interactions Through Decoder-Only Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.611175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.611175Z digest=sha256:82d4914dc585fde0b700265fb37488185cc676af46effa2cd113480fad46a4e6

Observation 6d269bcc-721f-4cba-ab9b-9803bb332f29 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Optimizing Vision-Language Interactions Through Decoder-Only Models InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.634043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.634043Z digest=sha256:aee9680b39f6cfde6d18152037ecfaa0ac9be6c0e04be71c2cc5fd373120299c

Observation 47267141-3a43-4ce5-9677-5de5bbb57551 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Optimizing Vision-Language Interactions Through Decoder-Only Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.628483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.628483Z digest=sha256:0c6fc0c614f2906ef168c6be92c8859591be309b98c57dfc913b9cb6e0e3fee0

Observation 2e08b26e-817c-4f41-9012-f37a34ff83ae · outbound

This paper cites Multimodal event transformer for i mage-guided story ending generation,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Multimodal event transformer for i mage-guided story ending generation,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.645686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.645686Z digest=sha256:5075d4e3419891e97922d11a95209a19f623b66fbf36c70087f0363415bc265e

Observation 290f8d76-452e-4d2d-bfc2-4da99cfc6534 · outbound

This paper cites TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens.

Optimizing Vision-Language Interactions Through Decoder-Only Models TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.639261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.639261Z digest=sha256:d85b6a61a93ac941c005f60c6a8fa8c6047370b3eef4571bc9312fbe407d5ddc

Observation aa2f60fb-8b6b-4693-8d29-832a263aaa79 · outbound

This paper cites Vilt: Vision-and-language t ransformer without convolution or region supervision,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Vilt: Vision-and-language t ransformer without convolution or region supervision,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:07.935301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:07.654957Z digest=sha256:df5d09b4df36c2c211a1f2f16f9eba38457ecdb80855bfe6f487bc71b83a7514

Observation 5b40c818-fa1b-4aaa-a51c-8228ed812c57 · outbound

This paper cites Style-aware contrastive learning for multi-styl e image caption- ing,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Style-aware contrastive learning for multi-styl e image caption- ing,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:07.953949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:41:07.650306Z digest=sha256:bcd2ed29d8c1ed0c1f45e9522e4fde13f74f657d5eeba8dac341b9f773d788c9

Observation d9efcc2f-8978-4d23-97a8-97fad21a7f40 · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

Optimizing Vision-Language Interactions Through Decoder-Only Models Unveiling Encoder-Free Vision-Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.412311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.412311Z digest=sha256:c200d91c85f359fda47a43b5161d05a8c201d31aee427349f2de9a3758840fa0

Pith citing papers

No inbound Pith citation observations are available.