Pith. sign in

Paper Citation Record · LEDGER

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model

As of 21 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2412.08307.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08307 v3

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:03:23.250097Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 30eef648-3d45-46c3-bbd2-169572f7c7c5 · outbound

This paper cites Mc-llava: Multi-concept personalized vision-language model.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Mc-llava: Multi-concept personalized vision-language model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.436256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.436256Z digest=sha256:bc7eb7d4e22838efd6a5ff7cc97acf5e69091a7a1cd4c5467c6e3217073e5d7c

Observation c8d76988-40d4-42ca-ba0b-88bfe23208ae · outbound

This paper cites POSIX: A Prompt Sensitivity Index For Large Language Models.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model POSIX: A Prompt Sensitivity Index For Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.469462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.469462Z digest=sha256:fdc0075f45fbd158af12e3a236705d3aec116aadf36e9aaf2940fc2aa911ef67

Observation 71b020b1-6cf7-446e-9f1e-99d3d252a245 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.525123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.525123Z digest=sha256:00ffebd3f6c73d51eccd3d8c202e1a81bd3e2596322ca46461a63bfea1dcb545

Observation 727b169a-47e8-49b8-964d-8bbdb9b750ea · outbound

This paper cites Sensitivity and Robustness of Large Language Models to Prompt Template in Japanese Text Classification Tasks.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Sensitivity and Robustness of Large Language Models to Prompt Template in Japanese Text Classification Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.539505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.539505Z digest=sha256:1bd670bd2c490627e434026c8e30b156e1759043158dbced7c0022576f03d9b8

Observation 74f1b90a-d998-4d21-9a53-24d00ff0a45f · outbound

This paper cites Demystifying Prompts in Language Models via Perplexity Estimation.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Demystifying Prompts in Language Models via Perplexity Estimation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.564900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.564900Z digest=sha256:4f6a184fa5b4847fddba3e8939699336eb9aff9faa60a740fba633e81eda746d

Observation f61515cc-88c9-4f50-8453-f6bab2cd4d0a · outbound

This paper cites The language of prompting: What linguistic properties make a prompt successful?.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model The language of prompting: What linguistic properties make a prompt successful?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:03:23.497784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:03:22.670661Z digest=sha256:55af763adac4a3442a6a3d52ae3a1019a066824c5bac1c165beed2911c14ad2a

Observation 38253995-78b9-43ca-8e9c-34e08fbf2b7f · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model LLaVA-OneVision: Easy Visual Task Transfer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.698234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.698234Z digest=sha256:3aa5b717c78eb04b03caa9f63491481620dd91ea32ac1619bd214b6ec711dfa5

Observation 1d211f89-446d-453c-bfb6-4df7a3a23281 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.726488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.726488Z digest=sha256:91e31eb84624069b83729c301e789700baae56edb0449bf1cdf88150f900f08a

Observation bd859e7b-4303-43e9-85be-3d014470063d · outbound

This paper cites Holistic Evaluation of Language Models.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Holistic Evaluation of Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.759469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.759469Z digest=sha256:39c6cdd63e716fe258dcd7c8919279b1ad14728ca38316f7c907f064ec4ae99c

Observation 892510f2-451f-46b1-8d14-8e4387de93e3 · outbound

This paper cites Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.802828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.802828Z digest=sha256:b54e86b6f786264b8bdc2801e699a64b032faab9cf1b44afed684146a273c4ab

Observation d192c475-6245-427a-939c-92dcdf3f2e87 · outbound

This paper cites Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.817035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.817035Z digest=sha256:22f99a3c8b9495646d6596bea568f903db6ba1b543136a54f99c4fcc7ad50d1d

Observation 6885cc57-b432-4ad9-8749-1deb724dae92 · outbound

This paper cites Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.832745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.832745Z digest=sha256:0195c677416e60830f3693fc2c2ccb29dd1c0aea1134f94553ad07e4b9569952

Observation e3995c41-b717-46a7-aee2-eeb732b4bd12 · outbound

This paper cites LLM as Dataset Analyst: Subpopulation Structure Discovery with Large Language Model.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model LLM as Dataset Analyst: Subpopulation Structure Discovery with Large Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.848861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.848861Z digest=sha256:09392a35a5b97c3ade82be9fe56d0a5a5defd43a3ae693d823d5858487aac4de

Observation cc2f8e6b-862c-4dd9-ad4d-c656dc517d3c · outbound

This paper cites m&m’s: A benchmark to evaluate tool-use for multi-step multi-modal tasks.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model m&m’s: A benchmark to evaluate tool-use for multi-step multi-modal tasks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:03:24.005992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:03:22.861857Z digest=sha256:d7f9852c21c20694d861747cc4402f481748f87bc7537b518849bae899c1131d

Observation 23712d74-aaa0-4c24-aca8-27258a256804 · outbound

This paper cites What makes chain-of- thought prompting effective? a counterfactual study.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model What makes chain-of- thought prompting effective? a counterfactual study

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:03:23.920333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:03:22.895898Z digest=sha256:ab995299f3b58c13668e6e04fbde463d0df3b6cfda8573d65ee18f0f8e064e89

Observation c7d9c0ca-fa84-4fcc-9c2c-24f50449ee61 · outbound

This paper cites X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.912520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.912520Z digest=sha256:5ba72235d902f6e15b4b70764b0abb9708b54eddc58df937aac8284435b1e42d

Observation 238224de-efcc-4bbb-b2f5-8e851586d489 · outbound

This paper cites Benchmarking Prompt Sensitivity in Large Language Models.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Benchmarking Prompt Sensitivity in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.960948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.960948Z digest=sha256:cb14a7db559c443a24ef231da295b028093cd6327e3bcfd06d2bbe145fa856cf

Observation 6628a6c9-03d8-42b3-a6ad-2eec363cc127 · outbound

This paper cites Mind Your Format: Towards Consistent Evaluation of In-Context Learning Improvements.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Mind Your Format: Towards Consistent Evaluation of In-Context Learning Improvements

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:23.086225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:23.086225Z digest=sha256:754e5d37e10088782bb469588a2b483af7a76f92b5c8070daf921e7231981247

Observation 0ace4f0e-b5ef-4ba9-9940-85c9ad3659f6 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:23.141039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:23.141039Z digest=sha256:6587d0080272bfcae6dbdb9a8ae5bc03ae8579f4e139207447cf03fbaa9458d8

Observation 60f3acee-f860-4bfc-954e-be4d3b88ffb0 · outbound

This paper cites Task Me Anything.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Task Me Anything

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:23.176924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:23.176924Z digest=sha256:d12a44d51767d7fe75dc13b52168f6ca6e69a9e9826849c2e829965c4a94d772

Observation 0008e30c-2ecc-4df8-85c1-79c3cc58b7b6 · outbound

This paper cites Large Language Models Are Not Robust Multiple Choice Selectors.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Large Language Models Are Not Robust Multiple Choice Selectors

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:23.211190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:23.211190Z digest=sha256:95926cca877a2ca0aa731ff95d46c3c80a06211b9f7b78892d2fae406f402f51

Observation bd150fba-9ea1-47ed-a48d-7b06b5411614 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:23.227142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:23.227142Z digest=sha256:12bdbce46bdf7a2aa8ef67b5c51863fb110b9dd6cf879fcc5df36d782ce811c9

Observation 44282cb8-b209-4021-acf4-b658edb810d5 · outbound

This paper cites ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:23.238535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:23.238535Z digest=sha256:10660ed561b5559fec915de5d7e889841343542f208209f33b356f4fca1a76b0

Observation 48e7bb19-ea13-4484-a973-706e8e2b0e26 · outbound

This paper cites Multitask Prompted Training Enables Zero-Shot Task Generalization.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Multitask Prompted Training Enables Zero-Shot Task Generalization

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:23.042512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:23.042512Z digest=sha256:fffaa3145aa124c1fc91c2ccd372369eeb92d6631b26d444c135818e1a989e0e

Observation f787eb87-748d-49bb-a266-288d1fe1fee5 · outbound

This paper cites What matters when building vision-language models?.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model What matters when building vision-language models?

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.645911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.645911Z digest=sha256:b68a81621a708768ba01a3f1b3b904441520a1b57bb2afbf5a63469f3c8ac53a

Observation 4215bf85-ded9-40f6-9e0c-ca1e6f48563f · outbound

This paper cites Scaling Laws for Neural Language Models.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Scaling Laws for Neural Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.621698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.621698Z digest=sha256:5dc296b2f3c1700749b724e456ce0687a8f775e082202d4752919cc05db52f76

Observation 4d106d51-ccd8-4f5b-bbf1-ab86482f5e93 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model LoRA: Low-Rank Adaptation of Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.595202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.595202Z digest=sha256:a976dd13123c2fa874e36942cf811f22cc57f9fc9e3b01788c14c43439c78a2e

Observation 5b33702f-5868-41dd-891b-12eae32a2e59 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.443790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.443790Z digest=sha256:45082b107f5c26f24477b6037e6aeef07559ebb7e6279ed80165c9b891ea8272

Observation d37fd50e-4050-4445-a0b4-86665cfa144d · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.501036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.501036Z digest=sha256:6fe217cd2d8bfa7c516db15677a4570bcd384e2ac4c3ae49445ff5d260d0b4f6

Observation c7344e08-9461-48d3-8f45-a6e95f959a77 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:23.010587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:23.010587Z digest=sha256:4c93f69ede8647cbf369ad8bf16772a5b2939cea7180e5f927e607c1d6f54775

Observation cc73d18a-015a-4dfb-8bb4-b8101aa3fadc · outbound

This paper cites an unresolved cited work.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model Unresolved cited work

Reference 3360

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:03:23.860969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:03:23.250097Z digest=sha256:28a802f4fffed4a1435342c6d92a1afcf15d087ff547e5661a6aa43b40e2929e

Pith citing papers

No inbound Pith citation observations are available.