Pith. sign in

Paper Citation Record · LEDGER

LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2503.15621.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.15621 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:39:52.261843Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T00:06:55.085465Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b210d165-7801-4b03-8c70-ea51aefc39d5 · inbound

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs cites this paper.

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.261843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.261843Z digest=sha256:b9cbde38fcb49d5aed53e4c3c891606c99bb77ce4889407c6084c0f9f10f59a6

Observation ddecc42b-11bc-4b97-b0cb-080f41e5267f · inbound

mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQA cites this paper.

mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQA LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:06:55.088914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T00:05:08.866244Z digest=sha256:dcead5c518382116955e26000df909b0bff93ed52b41011d07c70d80ec4a9fe7

Observation f550938a-05e9-49d9-9717-c33cde070ef7 · inbound

Towards integrated sensors for optimized OCT with undetected photons cites this paper.

Towards integrated sensors for optimized OCT with undetected photons LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T23:26:28.459673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:26:28.459673Z digest=sha256:8289929ae8f04a7fe0a24eb5e443cf8aa7f06eda17ddd38d8c6c4413cf928eee

Observation f4f18c87-847d-4d35-8a02-cdfb533decab · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:51.448384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:51.448384Z digest=sha256:a2da09046f61724772ab75ccd0d6dfde796f03680cf3b0456650beb5b180d8b2

Observation 65bd7cbc-95f0-43d3-8b85-aca1d21e623f · inbound

Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models cites this paper.

Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:20:59.258160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:07:13.367863Z digest=sha256:16621edd8e262c09fb2c5c427570ed7d6535d5dafc072d95cfb9803010498d54

Observation 16c8ce17-171f-40c2-9c69-059ea7b4bedd · inbound

Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance cites this paper.

Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:07.336220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:27:54.092750Z digest=sha256:a0043e12a0f64b889eca5e793cd386373edc991ade822d144cacd1e2c9478749