Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:49:34.963022Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 4 inbound Pith citation observations for arXiv:2507.09876.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:49:34.963022Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T01:52:44.785582Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T12:46:56.833017Z
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7cc29dba-3559-498c-aca7-61e5c70afd2c · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27dccd33-300a-42ae-bdef-f6fc5854b9cf · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Hadzic, Taran Kota, Jimming He, Cristobal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Li Fei-Fei
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 015d76b4-6404-43c0-8a9a-186d849caf1f · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f229cfb-9c15-459d-ba5f-197b93d0830e · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27225dd8-3d9e-4ee7-ba66-f364c93899cb · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b51a282-7bef-46aa-bf45-3d01fdc5038b · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models AI4Research: A Survey of Artificial Intelligence for Scientific Research
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 381c4fb1-79fd-43e4-874f-f9ace96b357c · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cbc93df-fee0-4c19-be81-59e670680f03 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d461dc0-9426-464a-83f7-fcb862011ed3 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a1f1955-fb2e-4578-932a-2a72e29473a0 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cdb2164e-c249-4acf-8c87-f125a9f45085 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c9ba49d-6aa0-4872-a967-faf832a01226 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2612be1-665c-4f42-aef0-71123dcd7b7c · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc9b25a2-4771-49c2-9ba0-ada62172607b · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 95d610e5-de64-4d0d-9cfd-0afd5792bf15 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47f20178-e7b6-4f49-b8e5-6375af45dfaf · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40a9f754-3d02-443c-a189-4b4386002a73 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 34ee85ce-7b5b-4e31-ad35-2200692b9b20 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2608cca2-ea4e-443e-af2f-9c56d48093d5 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models CoS: Chain-of-Shot Prompting for Long Video Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c83bbcbc-7685-4a7b-a003-7427cf56afec · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a6a185a-9cf7-4d98-8d98-077db591bbc2 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2931641-5298-47e4-9ec9-3fe25df594a7 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf46c7f1-1d0f-44e2-aeb7-1e2d646f3f9a · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 705c3741-17e4-41b5-bba2-d05810d930e0 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99ae4b9e-a7b2-4628-8444-243c9f642b85 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f5a3dfb-fe81-4984-a699-8009d92c7ac0 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Large Language Models Meet NLP: A Survey
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 543ab5dc-eed6-4180-9e03-722476cbd432 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 417c138a-f82d-4e25-8533-f5fecc693e0f · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaf1e309-0123-40fd-9642-f6257d239dee · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdb59fb9-855e-4629-98b2-c2e8e7362bd0 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b118f7e-1b0c-483d-bf9e-cf69301697dd · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8b75389-fff8-4870-8dad-a03a8435da70 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4367e60-f478-420a-ac14-fdb8ded6754f · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e94e0bfa-8ec6-4cfa-aa30-3dd8f9994c8e · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96400931-ec18-4727-a50d-ab57e09dac6f · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4039a2da-2b2a-458d-952c-ce10cad2ac97 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7b84ee6-057a-4995-bdb6-9daf7fb89eb0 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 07208b62-5b89-4fe8-a7d5-d9468087a54e · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 725d5f29-1d0b-4b20-ad2e-647f8f1ac83f · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa414a16-3ddf-4612-b5aa-f74256674037 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4021bd4b-b315-4079-99d1-370aadc0c1c5 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75444443-fc2f-4161-a82e-a565c3d2ca3c · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81bbbdab-83a1-4ccd-b08e-068c332d6e55 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models In Proceedings of the 32nd ACM International Conference on Multimedia
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f5c01e62-6511-4ff8-a8b5-38d91ddb0294 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 738d0cee-20c3-4b20-9962-7b8c1bdad03c · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Beyond Chain-of-Thought: A Survey of Chain-of-X Paradigms for LLMs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de0f54f3-7733-48c3-a7ca-14a2344bf4da · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dce07eb-d539-47f4-b7e2-3ff2e92a9a68 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da72bc3b-1d72-4d21-b00e-4fb3eae1daec · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb1ad447-fb3d-417a-bbda-ba92963928ef · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Qwen2.5 Technical Report
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1269154-5f01-4e48-9464-a743702026bd · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5dfb38a4-4956-47d4-9875-0bea4f982ff9 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e3c5ab6-bafd-4e4b-bf04-bdd6ff960e1a · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 880af460-c720-488d-9016-1775980c506b · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f6682e-d6f3-4804-97bc-14acf2efeb61 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models A Survey of Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e905d582-7990-4c7a-a245-8611b3d2a1ed · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models CCHall: A Novel Benchmark for Joint Cross-Lingual and Cross-Modal Hallucinations Detection in Large Language Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd4dff17-aeac-4d24-8f70-b2c93b50ac69 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models MLVU: Benchmarking Multi-task Long Video Understanding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5048f173-f12d-46c6-b396-352c5b0f0fb0 · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71857a80-a104-4735-aadb-a2a0ed40550d · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88264853-d586-461a-bbb6-a98f2bc7408b · outbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models In European Conference on Computer Vision
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 558a108b-cffe-40fe-a71d-db26eebee2c9 · inbound
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51a2c363-4137-4637-9b3a-ed92ff0003fd · inbound
Act2See: Emergent Active Visual Perception for Video Reasoning ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05811938-693a-4663-9912-8d5666776b84 · inbound
OProver: A Unified Framework for Agentic Formal Theorem Proving ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aed15bde-fead-4786-81d4-3742e7779b81 · inbound
VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.