Pith. sign in

Paper Citation Record · LEDGER

How Important are Videos for Training Video LLMs?

As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2506.06928.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06928 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:51:07.923531Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:40:00.723218Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65c748a2-cd0b-461a-97a9-c9cd92477f4f · outbound

This paper cites Qwen2.5-VL Technical Report.

How Important are Videos for Training Video LLMs? Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.773495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.773495Z digest=sha256:e651e93a7aee712782b6e0bad270fa5114531a74bca752813f9882df05a9ab8d

Observation 285cf941-cb86-40de-8a0d-f2cc99af891e · outbound

This paper cites ShareGPT4Video: Improving Video Under- standing and Generation with Better Captions.

How Important are Videos for Training Video LLMs? ShareGPT4Video: Improving Video Under- standing and Generation with Better Captions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.459611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.779399Z digest=sha256:a0376a3cc30658181f5164be88e6d0998be470879e896ff95544cc3cfd1b6dca

Observation ea581482-7c3a-42a9-a25d-ee24df90f017 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

How Important are Videos for Training Video LLMs? VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.784472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.784472Z digest=sha256:008a458e9b1bb043fa0d030193eef3ef31969d20bdc0566dccdf8e7edcbd32df

Observation 03cff194-f5a0-4074-90c0-a2701db95f6b · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

How Important are Videos for Training Video LLMs? Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.790019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.790019Z digest=sha256:4ba97b0c0395bc7982dd64aa197169915a56b5436fc2df18762324f6eea0ad64

Observation 4f6aa8c0-6d3b-4c63-9c20-f096c821bc1d · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

How Important are Videos for Training Video LLMs? Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.795669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.795669Z digest=sha256:75ab6e57d777d49dbdcd6f95fc2c47124fcf81b877255feaf321e1179d7a94bb

Observation 1c2a4e8b-65ba-4337-a695-abd2ff200285 · outbound

This paper cites The Llama 3 Herd of Models.

How Important are Videos for Training Video LLMs? The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.800928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.800928Z digest=sha256:fa9708a6be49a9dd381619239bc18c07a9f886f49451c50edce01b22a331f1ad

Observation bf156721-0a82-4991-b65e-8aa7f00a3508 · outbound

This paper cites MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding.

How Important are Videos for Training Video LLMs? MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.443799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.806446Z digest=sha256:49f9f750b4d556bcbe876f426b0f526e1a13b354e1bda101e416a295e04cce1e

Observation 063e0f5c-e995-4454-a493-83357a8972d4 · outbound

This paper cites LoRA: Low- Rank Adaptation of Large Language Models.

How Important are Videos for Training Video LLMs? LoRA: Low- Rank Adaptation of Large Language Models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.427508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.811592Z digest=sha256:efc7f705376bfbcc542d18595e0e26c84ee98026db8b92d7789fe679115c2efc

Observation ed31497f-5f07-422d-869b-cc8a4a3fd827 · outbound

This paper cites TGIF-QA: Toward Spatio-Temporal Reason- ing in Visual Question Answering.

How Important are Videos for Training Video LLMs? TGIF-QA: Toward Spatio-Temporal Reason- ing in Visual Question Answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.408349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.817057Z digest=sha256:0679e0f39108109e0e8bafdf561b41d1f34e0f2603e433c7a4548304ad3785ef

Observation 09ad63b2-12d0-452e-b6c1-6bcae633622a · outbound

This paper cites CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning.

How Important are Videos for Training Video LLMs? CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.389004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.822018Z digest=sha256:b66de9b61694d64b3c86f5ccd986139a3eba50b0f515f7921b9928a9e2d73c29

Observation 053fb2fb-be20-41d5-ad0e-0dc963146e8b · outbound

This paper cites JUWELS Cluster and Booster: Exascale Pathfinder with Modular Supercomputing Architecture at Juelich Supercomputing Centre.

How Important are Videos for Training Video LLMs? JUWELS Cluster and Booster: Exascale Pathfinder with Modular Supercomputing Architecture at Juelich Supercomputing Centre

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.373009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.827043Z digest=sha256:cd3f609f6a6389c98e81ec4a67952d2a0dc82798e15324b9d66e840e5dc4f49b

Observation a8b5ff99-15ec-41f0-8177-5e10f582f342 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

How Important are Videos for Training Video LLMs? LLaVA-OneVision: Easy Visual Task Transfer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.831670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.831670Z digest=sha256:8509debdd1360dc909b7cca02171e571fbae1d5cc91fe7efb637517328b47992

Observation 39381d6c-50ee-468e-a4a9-55417ca9e93c · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Under- standing Benchmark.

How Important are Videos for Training Video LLMs? MVBench: A Comprehensive Multi-modal Video Under- standing Benchmark

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.357663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.836882Z digest=sha256:1305a3793a61435d556f180d74b5a6a7fd46c7aefbba1aa1204d9e11053b94a7

Observation ad71e2fe-5895-4c17-814c-8f8992cfa555 · outbound

This paper cites Temporal Preference Optimization for Long-Form Video Understanding.

How Important are Videos for Training Video LLMs? Temporal Preference Optimization for Long-Form Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.841588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.841588Z digest=sha256:ce7dff9317053f9932ecd70665014fb527839e02bc040fcbfdbc7dad248c99c5

Observation 6e1e2824-7908-4c1d-9206-0cf2f47da698 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

How Important are Videos for Training Video LLMs? LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.342865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.846524Z digest=sha256:1d3a5181e98ebbef931a85031e26094ab1a7e8ff7368cfd4a1a6ab090024bbe7

Observation 9f8edbfe-10aa-4d2c-ae75-c677098d699e · outbound

This paper cites Video-LLaV A: Learning United Visual Representation by Alignment Before Projection.

How Important are Videos for Training Video LLMs? Video-LLaV A: Learning United Visual Representation by Alignment Before Projection

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.327932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.851065Z digest=sha256:a0252536ee206ff260203e39c24ab4b5dcf211c2696cf90a7923ea4d3ebb3e99

Observation 5b7ce0b0-5171-461d-a9b3-8f1314059ba4 · outbound

This paper cites Microsoft COCO: Common Objects in Context.

How Important are Videos for Training Video LLMs? Microsoft COCO: Common Objects in Context

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.312225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.855756Z digest=sha256:a35339e8e731896e35db1c8016c964bf9c7c401d83d598c87f414bcb90a1c44d

Observation 9ebd11f4-cb7d-4fc7-9870-6582a89258f8 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

How Important are Videos for Training Video LLMs? Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.860377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.860377Z digest=sha256:d5e89fb899d5ef3a5d24ba07bcc017e929c0bea8f8b84669c10fa0bc5fca1dca

Observation bf28ca4d-4cfe-4ed5-8dd5-db851bd7a74d · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Under- standing via Large Vision and Language Models.

How Important are Videos for Training Video LLMs? Video-ChatGPT: Towards Detailed Video Under- standing via Large Vision and Language Models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.296144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.865389Z digest=sha256:9ad827ce346dab6e72d760e3f777c85095b7a1c7f3749bedccc6241635be8908

Observation d9d66231-4270-478c-bf57-9fa8221d2497 · outbound

This paper cites an unresolved cited work.

How Important are Videos for Training Video LLMs? Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:51:08.279071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.869984Z digest=sha256:c47a91deaa249ec4bc7f875355329087f5f754403f083045ee36e4e5246db9f8

Observation f86bf564-bf9b-4076-b424-f45354ce05a1 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Super- vision.

How Important are Videos for Training Video LLMs? Learning Transferable Visual Models From Natural Language Super- vision

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.262254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.874688Z digest=sha256:ee45d719b51a4ed219d9c9a9abd031ce3885aa76c5fd974aacd55c7ac27cb105

Observation 6dce3466-df81-459c-a18d-87100f0d2f40 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

How Important are Videos for Training Video LLMs? LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.879150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.879150Z digest=sha256:22fb822df1fc7bb9309054593e585a94434b8893b8d99f89378e5826febc27ca

Observation 0cb707a5-6917-4e9f-b00d-1c255f096643 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

How Important are Videos for Training Video LLMs? Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.884760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.884760Z digest=sha256:6d9c3d722312e00297d2d03b1e8285d13baee77d945d99b6fe9cf85a66e36fe5

Observation c3ee944a-08d6-479e-a3b3-22b187f7925e · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Inter- leaved Video-Language Understanding.

How Important are Videos for Training Video LLMs? LongVideoBench: A Benchmark for Long-context Inter- leaved Video-Language Understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.244690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.889776Z digest=sha256:b2e4a0ebeada655573d738073a55b43f3156c7a1e78efc7edf8b2fe21353c0d5

Observation 4e0dc231-9de6-4f15-8f1e-0112554d8688 · outbound

This paper cites MSR-VTT: A Large Video Description Dataset for Bridging Video and Language.

How Important are Videos for Training Video LLMs? MSR-VTT: A Large Video Description Dataset for Bridging Video and Language

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.228268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.895050Z digest=sha256:52b9d768facd08cb522bac8853b70b6f10f63d6f15ea94603c2f0347cc6a133c

Observation 897dc92b-73ff-4532-9744-3ca2e42c1f23 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

How Important are Videos for Training Video LLMs? PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.899664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.899664Z digest=sha256:3fa25b908e66c04a1e6fad268d561d6daf3389d9e182cb54bfc5d278c3db46eb

Observation 87ec161b-a15b-4e8c-afd8-4be53ef31414 · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

How Important are Videos for Training Video LLMs? CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.211642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.904430Z digest=sha256:3794cdfc38966bf30b1d38c4dcaba1c5c6abfe533044409d9cf0f1d137b57781

Observation ff108d37-dfb5-4f80-8fa6-b5a81e09e326 · outbound

This paper cites ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answer- ing.

How Important are Videos for Training Video LLMs? ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answer- ing

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.195305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.909084Z digest=sha256:6d37c647bab1869918d7dfa62ee5a592d2dd207a1550d5bebfc90073e1538590

Observation c96d5b97-b387-450b-a5f1-13dc333d2a9f · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

How Important are Videos for Training Video LLMs? Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.913992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.913992Z digest=sha256:6e716424ed1e1d20cd411438bc870eae0167502c430d5108dabe65c64720b1a0

Observation 5faf4ef4-d8e6-4340-b5db-458784a5186f · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

How Important are Videos for Training Video LLMs? Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:51:08.177820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:51:07.918987Z digest=sha256:cfd95248435a861f35adc13f1956bc5e9bdedf7317683444f9f3c230be57df3a

Observation b6f15c9d-18d7-4e73-9bcc-4a2501bf2bc6 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

How Important are Videos for Training Video LLMs? LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.923531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.923531Z digest=sha256:38b00f30d53b8801db4d48140fe01f65d78b991c8f3123a31fe607b824fdb9ea

Pith citing papers

Observation 8c7647d5-6f3e-4110-8c64-34e90ee76aa9 · inbound

Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks cites this paper.

Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks How Important are Videos for Training Video LLMs?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:00.723218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:00.723218Z digest=sha256:fb3d6d95b4624b38865c24cedf259b0f626b47f6239be6e54bed7ce251f0baf7