Pith. sign in

Paper Citation Record · LEDGER

LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2311.17043.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.17043 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:33:19.326679Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T19:35:01.039481Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 973d328e-1357-49cc-b90e-1af090db8bb9 · inbound

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation cites this paper.

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:55:20.473273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T04:55:20.362512Z digest=sha256:c1346bf04c7fdc0059c3efbe4af2e710cd3ca8abc80d452b789082bec56c1147

Observation 73a3776b-4a93-41a3-a408-726ba5d7b8c5 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:16.731922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:3093f142da20201c6fab8b391c7468ba86232501ea0158d4d1563eec68ab12a3

Observation 081ced19-1148-469e-8efd-826183496418 · inbound

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models cites this paper.

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:44:47.497607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T07:44:47.355960Z digest=sha256:d21878bf74b9c3b230808db45d034f326dabe7d90d6c3407ca6d1423fc349619

Observation 61fb005d-633c-49e3-a414-192694e8cc4c · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:21:57.974849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:177b4b99ef2c70b3da0be321cbc1406f6214ae94bd7cdaa6f4f80b3d0201fb99

Observation d26af7f3-d15e-46d6-8cfc-cf9220c4f338 · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.389857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:64e1b29f08c271f16b0d7e132dcf0a2659a5289d605b2225e97adcd7ba17cf77

Observation 31eeb163-b9d7-451e-ad84-70b55c64fd6f · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.195081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:125a5571f7748ab5db32245f636d05cabbde006a5b51bfed2b26b3e7e5792cf3

Observation 7b089da6-9c35-461c-bc2c-368e96b13c4d · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.654562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:0a56a380afb5a76a16f2209d675ace86fb95d1244e2442463d7d9d2306005649

Observation 075a7d8a-4684-4e04-b3e0-ef9b26d2067e · inbound

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models cites this paper.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:54.037741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:a9d4a1ef7ff57c6220f0c81e7a25df59f5dea1783d36d310a2445df027c82bf0

Observation fc8b33c8-ddc1-4220-9ece-ac46ead134ee · inbound

What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction cites this paper.

What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:53:33.157114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T22:51:08.753650Z digest=sha256:73117cae97534c9b39594bd59769888bc79e787ae82ae8f17eafdecd3a472246

Observation 170e0518-604d-4cd6-a103-28c2ad1cb0fb · inbound

LongVILA: Scaling Long-Context Visual Language Models for Long Videos cites this paper.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:51:25.511626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:4fcce7ef3febf03df2f0900249ddc4b6010a9b6495f1b183719922cd89aab6aa

Observation 81baf91d-9190-46e9-a25f-7aca2b3d8d6b · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 193

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:32.978337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:306ce8e8fe77c2336c548cbe4177bb3d8b8b27e5b151843fc043f7a8253936b9

Observation e6fb5e39-367d-4e45-a0c8-deebae6711f9 · inbound

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks cites this paper.

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:51:36.333783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T19:51:36.137985Z digest=sha256:70fa82a8848ca19fd14eb0419bcc609d0583984ad4152f86b7e58ca60384f9a0

Observation fa43ef72-8084-4d98-b812-6fccee3b5343 · inbound

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model cites this paper.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.326679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.326679Z digest=sha256:87f2d7357dba9832fcf9ca224bfab58e09073ab543a3602fa2272415451ad2ca

Observation f9bd112b-a33b-408d-b03c-4849703a8fb1 · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.193817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.193817Z digest=sha256:8ac869cb587ff6b8789ab8d7febd75fe96a13600f003e5bdc73eb0ed10e72145

Observation 3de1bfdc-6e02-41d3-98f1-9a17c5964c27 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.514961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.514961Z digest=sha256:aa982e7a158c0bc3d3a6858d8c1ac16677767596870ad44a806fa2a7bc7c3ab6

Observation 2efa0987-9240-45d2-aa06-4a09a803cc06 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.144931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.144931Z digest=sha256:302a16824f7b3033c0df17ba58ebe4fd85512066f5b802f5310f4f66f280ed45

Observation 00d64863-bf51-4eca-aa1e-a18a39129839 · inbound

SV3.3B: A Sports Video Understanding Model for Action Recognition cites this paper.

SV3.3B: A Sports Video Understanding Model for Action Recognition LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.740370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.740370Z digest=sha256:356258d0f6e8e5cbd3e0ad08bfc9b7888772e6b2b85de14a6b2f76b9c7a50312

Observation be12beba-f4bf-4e8c-96dd-2118725ef2ad · inbound

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs cites this paper.

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.230172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T03:18:11.993413Z digest=sha256:29f3c25671417dfae8889fe253917af63335234d69006b93ac31f22942eb3e39

Observation f98a0e98-a01c-4b5a-a1a9-de828b50673f · inbound

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding cites this paper.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:01.099735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:01.099735Z digest=sha256:3c356a8d1a82a066caf585c9b3b1f02174863e698c57dd8166cac1029d703510

Observation 68cc46ca-92bd-4dd7-9f09-73c833ef94d6 · inbound

Video Reasoning without Training cites this paper.

Video Reasoning without Training LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:08.173803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:12:08.173803Z digest=sha256:26cb839410039cc4b216aa7b1f56aa18dadb6d372c764f7f9ffbe646b2f30791

Observation be4b702d-4fc2-4de9-bbd3-c0e28db84eaa · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.347146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:c7a52ccdb8c5df57f69a9444ba7b859aa9005deb73f674f16981bd8e707b7376

Observation dd36a994-d8b2-4431-b287-a07b92f577df · inbound

TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation cites this paper.

TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T21:37:47.904757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T21:36:43.971666Z digest=sha256:4bdba58694ad5ee9fa6de756c27fe240a0169c7d2e257a062e118349bb913faf

Observation bf2a9c55-f062-401e-9692-ed22c84b8528 · inbound

TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation cites this paper.

TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T19:35:01.040957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T19:33:24.558488Z digest=sha256:f14968e9cc8eece4fb81d0618bb039de68a4a940144572cb215b9443cc156842

Observation bc6cffe9-4ff9-4f63-b216-8efe10b6a225 · inbound

AffectVerse: Emotional World Models for Multimodal Affective Computing cites this paper.

AffectVerse: Emotional World Models for Multimodal Affective Computing LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:48:05.768202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T06:46:33.612905Z digest=sha256:0f1cfd109b7b8cea27e0e913f676ed9284102f5af6e6f9bb75f3936c807e97ac

Observation 35d8a661-3f92-4a48-94f2-fc65384c5157 · inbound

Latent Visual Cache for Video Reasoning cites this paper.

Latent Visual Cache for Video Reasoning LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:6153d9b57b7bfb78b2b798f468c9fdfb0ad5d83215d1d3230fb9a216977cead3

Observation fa248f94-9bc2-4a6a-8b45-563b728ff2ec · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 188

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:7eaabe1f6bb1ebcf89dd30f53ef13c2996d6f573d89d4262fedce06a0fa11cd6

Observation 6da9c10c-6546-4c8f-b1d6-032321ab34ad · inbound

QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding cites this paper.

QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T17:18:41.284513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:18:41.284513Z digest=sha256:b44ee30bfe374e4651ede6bad63e2a2df56f0e9601d2684051fa886a60cb1b9e

Observation 4228e8d1-8124-4ab5-b3d3-f7c6d59913a9 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.729612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.729612Z digest=sha256:3d6bd2ebe5ed9111ed7cda8dbbd5a4b267ec9dc8e69394d0efe0dde69db886c9