Pith. sign in

Paper Citation Record · LEDGER

LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2503.15621.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.15621 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:40:42.521585Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T00:06:55.085465Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 93706ca6-5114-46ac-8f68-e5e62870d8b9 · inbound

MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models cites this paper.

MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:40:42.521585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:40:42.521585Z digest=sha256:e3e0ca4d3aade5ef90ee6d106de9d908fa1c32855e0ad0aaefcb514ec3b9d9d0

Observation f8133719-f5f7-4b14-8706-d58df949284d · inbound

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering cites this paper.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.705368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.705368Z digest=sha256:9959ff3c37cd428d946aa1353d0c0d4d7f59c702b9edd5301245b650533b9419

Observation b210d165-7801-4b03-8c70-ea51aefc39d5 · inbound

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs cites this paper.

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.261843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.261843Z digest=sha256:5931c9af536405c212b5ab5c12a2c8047aabe718eb85e09877d9e2732e545696

Observation ddecc42b-11bc-4b97-b0cb-080f41e5267f · inbound

mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQA cites this paper.

mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQA LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:06:55.088914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T00:05:08.866244Z digest=sha256:2e212a3a46344d1f9496f273158dd391502a05028ab490c45055f68263265a9e

Observation f550938a-05e9-49d9-9717-c33cde070ef7 · inbound

Towards integrated sensors for optimized OCT with undetected photons cites this paper.

Towards integrated sensors for optimized OCT with undetected photons LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T23:26:28.459673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:26:28.459673Z digest=sha256:7c875109a51f2855288946adc9845a81d8edfb436891ff7db3ae6dcb555e19a4

Observation f4f18c87-847d-4d35-8a02-cdfb533decab · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:51.448384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:51.448384Z digest=sha256:a485df0dfaeb51259e366649b4f88e2dea687393b58b3d52fb5760167a42eaea

Observation 65bd7cbc-95f0-43d3-8b85-aca1d21e623f · inbound

Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models cites this paper.

Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:20:59.258160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T16:07:13.367863Z digest=sha256:fbcce1cdcd2f0567fd9fb684121691369e33160cdd0bcc87635636aa037c86d9

Observation 16c8ce17-171f-40c2-9c69-059ea7b4bedd · inbound

Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance cites this paper.

Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:07.336220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T14:27:54.092750Z digest=sha256:b3e8003a8598aed56403a0f3e809b6faaea8d711d4307acb03f35684fe547c54

Observation 729e283c-fe58-406d-a371-03d8e94288e4 · inbound

Multimodal Model Diffing for Feature Discovery and Control cites this paper.

Multimodal Model Diffing for Feature Discovery and Control LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.798668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.798668Z digest=sha256:53a27bc8cc07f15259954477acd4639a480734d3315e5181f62aec74a64d031b