Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2311.17043.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:33:19.326679Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T19:35:01.039481Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 973d328e-1357-49cc-b90e-1af090db8bb9 · inbound
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 73a3776b-4a93-41a3-a408-726ba5d7b8c5 · inbound
TempCompass: Do Video LLMs Really Understand Videos? LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 081ced19-1148-469e-8efd-826183496418 · inbound
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 61fb005d-633c-49e3-a414-192694e8cc4c · inbound
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d26af7f3-d15e-46d6-8cfc-cf9220c4f338 · inbound
MLVU: Benchmarking Multi-task Long Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 31eeb163-b9d7-451e-ad84-70b55c64fd6f · inbound
LVBench: An Extreme Long Video Understanding Benchmark LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7b089da6-9c35-461c-bc2c-368e96b13c4d · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 075a7d8a-4684-4e04-b3e0-ef9b26d2067e · inbound
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fc8b33c8-ddc1-4220-9ece-ac46ead134ee · inbound
What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 170e0518-604d-4cd6-a103-28c2ad1cb0fb · inbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 81baf91d-9190-46e9-a25f-7aca2b3d8d6b · inbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 193
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e6fb5e39-367d-4e45-a0c8-deebae6711f9 · inbound
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fa43ef72-8084-4d98-b812-6fccee3b5343 · inbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9bd112b-a33b-408d-b03c-4849703a8fb1 · inbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3de1bfdc-6e02-41d3-98f1-9a17c5964c27 · inbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2efa0987-9240-45d2-aa06-4a09a803cc06 · inbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00d64863-bf51-4eca-aa1e-a18a39129839 · inbound
SV3.3B: A Sports Video Understanding Model for Action Recognition LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be12beba-f4bf-4e8c-96dd-2118725ef2ad · inbound
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f98a0e98-a01c-4b5a-a1a9-de828b50673f · inbound
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68cc46ca-92bd-4dd7-9f09-73c833ef94d6 · inbound
Video Reasoning without Training LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be4b702d-4fc2-4de9-bbd3-c0e28db84eaa · inbound
Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dd36a994-d8b2-4431-b287-a07b92f577df · inbound
TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bf2a9c55-f062-401e-9692-ed22c84b8528 · inbound
TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bc6cffe9-4ff9-4f63-b216-8efe10b6a225 · inbound
AffectVerse: Emotional World Models for Multimodal Affective Computing LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 35d8a661-3f92-4a48-94f2-fc65384c5157 · inbound
Latent Visual Cache for Video Reasoning LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa248f94-9bc2-4a6a-8b45-563b728ff2ec · inbound
MentalThink: Shaping Thoughts in Mental SVG World LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 188
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6da9c10c-6546-4c8f-b1d6-032321ab34ad · inbound
QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4228e8d1-8124-4ab5-b3d3-f7c6d59913a9 · inbound
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.