Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T17:43:03.400998Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2411.12355.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T17:43:03.400998Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
80 of 80 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2192909b-a3ec-40ac-aa7f-5a5b0ac5a6cd · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ccf63729-4ef7-4c31-9099-15657401fcf3 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Textvqa: Towards understanding of visible and invisible text in images
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a7df40f1-6a21-465e-8c84-95e39da5afaa · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae255e3c-586c-4dc1-aa72-d127c3ac2d32 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Frozen in time: A joint video and image encoder for end-to- end retrieval
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 036feb4b-a3b7-4b69-93c4-d10d7eba5f3c · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Frozen in time: A joint video and image encoder for end-to- end retrieval
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 47c3774e-a343-4e4e-9a7a-3209d40cd88d · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Learning with Differentiable Perturbed Optimizers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 169e7c30-2c2d-47b9-b16c-25b842bf533a · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Activitynet: A large-scale video bench- mark for human activity understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e2174da7-e12b-41df-aeb8-06e665057820 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Chen and William B
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8c0bab9f-0d08-4aef-b355-b32aadb75f55 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c0a61f2-f1a3-46bd-9800-de413e0a100f · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Gonzalez, Ion Stoica, and Eric P
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53588708-e7e8-44db-86f2-a24f74c2d2c5 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Palm: Scaling language modeling with pathways
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a94e7e06-e704-43d2-8ba4-b05789559123 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8923c3c4-b5a4-4d60-b4ec-53ed7bae0b2e · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding EgoQA: Egocentric question an- swering
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5178ef37-024c-4463-8b7d-feb5f9a4adf1 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Study on density peaks clustering based on k-nearest neighbors and principal component analysis
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 114aef3b-3e26-4775-92a2-779f5607bec3 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding EV A: exploring the limits of masked visual representation learning at scale
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7a073534-89e0-4cd5-8319-8f235bfaaaef · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2b2efb6a-fb1c-44e9-ad25-53a1c8e3cfd0 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Semantic-aware modular capsule routing for visual question answering
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 967e23a5-a606-4c25-8945-23b63da30962 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MA-LMM: memory-augmented large multimodal model for long-term video understanding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5de07298-2e8b-4d31-94ad-f6ed51167c0a · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6d4b9996-60c8-438c-b156-f73174c9c75e · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Activitynet: A large-scale video bench- mark for human activity understanding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3d217d76-55f0-4919-988e-a7158c81f2f2 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 02a8aba9-ed37-42e7-a542-a2bc61b21d20 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding VTimeLLM: Empower LLM to Grasp Video Moments
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4a3e019-fd4d-42d8-8cec-21765358ad1e · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Ng, Hongqiang Rong, and Zichen Li
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1c862dea-09fb-40c2-8280-c60f3386f8ad · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Gqa: A new dataset for real-world visual reasoning and compositional questions
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a63abfb8-567c-49a0-aabe-73c348fdd3a4 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2611f948-041c-4fcb-8df2-73f54e44b3ed · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Scaling Laws for Neural Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b32a324-e24f-4f7b-ac0d-ebe87d4d61ba · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d1aedadd-e8e9-4b7c-bc4b-b87f8f061e44 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Kingma and Jimmy Ba
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1358aabb-24c5-4dfd-84b4-abf6534493aa · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fe1aeb9e-49b2-4b88-8ada-5996889d545e · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8dd26375-64c0-4b7c-87be-1d82ecc6d9e3 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b36e757-7f2d-47c3-b746-91d1986490f7 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d816dd29-779c-4244-9cf9-18ca9a8829ea · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Scienceqa: A new dataset for science question answering
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation afc0463e-dede-48a9-9cd3-afd740fe1de1 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Learning dynamic routing for semantic segmentation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eb7e9455-6eaf-461f-8f19-563a7e6d1ab0 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66d3c174-85f9-447b-a60a-1f540dc777cb · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Video-llava: Learning united visual representa- tion by alignment before projection
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c1e97409-d45b-42bf-98a6-7a84d74224b8 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90fcb0c1-438f-463f-8bb4-6fd700a3c9d0 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Vila: On pre-training for visual language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 79bff67f-cb71-4e27-9c1b-2d75451d3a95 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Lawrence Zitnick
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 25d71ced-da19-4f2f-a516-afa49035b62f · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Improved baselines with visual instruction tuning, 2023
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de0d25fb-e774-4205-9732-0b5a640fbf72 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60a8e0cb-1022-4047-a7ed-825647eab5e4 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1593324f-320e-4af4-9cb9-cfc8b1ed7bc7 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding ST-LLM: Large Language Models Are Effective Temporal Learners
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2af014f9-11db-4cdd-aa50-d2a8f3d2c860 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ffbb597-19fb-4961-93fd-8e4f31a16a59 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8a8055a-73b3-44f0-91dd-8e70d97aa0ee · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Some methods for classification and anal- ysis of multivariate observations
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1f1405bd-bdfa-4a08-b6e1-179d714a455f · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Yuille, and Kevin Murphy
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fbd14a40-e224-4093-8831-e1162e9cf647 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Ocr-vqa: Visual question answering by reading text in images
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba16ab19-d20a-4e06-9572-00797ee7c6ab · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Webvidqa: A large-scale dataset for video question answering
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6ea3099d-e5d3-48e4-be36-b99bf90505b1 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Introducing chatgpt
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b9e6b080-9048-43f3-8b18-406637cedd42 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding GPT-4o system card, 2024
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6ecae718-bdbf-4511-a75c-0ff1b842447e · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Learning transferable visual models from natural language supervision
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b2a3c4cd-944a-46b1-a257-bab491b451bf · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Improving language understanding by gener- ative pre-training
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 80ab209e-d25d-468f-b4f5-f2c99e816a88 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 849ce177-89ab-4c57-84b2-263947569873 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c8db052f-139d-4721-92db-375d350333a0 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Massof Sarah L
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e74eb7fb-e81a-4a1a-aedb-f5c5f9d1a06d · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding A-okvqa: A bench- mark for visual question answering using world knowledge
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 35c00115-a116-4de1-9640-eabf185e24a1 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 04cbea46-a537-4ccb-a7c1-a20ce0943b2b · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Textcaps: a dataset for image captioning with reading comprehension
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 07fcd0e9-4a9d-48f4-913e-16a82ba77ff9 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 244acf83-61ea-4a6b-9dd9-53248aa8ebca · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Moviechat: From dense token to sparse memory for long video understanding
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation adc1f000-41c0-4e15-beea-8be646f352d0 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f27cfd0f-7d32-4381-9795-91a3214f24e6 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding LLaMA: Open and Efficient Foundation Language Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43a75b67-c0a8-435b-9aac-d5c43325d460 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Ocrvqa: A new dataset for optical character recognition in visual question answering
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7e892c22-933f-4318-965b-dc19ed6cf8a9 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding LVBench: An Extreme Long Video Understanding Benchmark
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97f740c3-73c6-4c9e-8958-67f8bdfdf36e · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd29c8d1-fd60-4a56-b0ad-e2dc3962211b · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Videollamb: Long video understanding with recurrent mem- ory bridges
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cc9702e0-bbdb-40bd-a4b7-33a4e5b83220 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding FreeVA: Offline MLLM as Training-Free Video Assistant
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d789eab-03a8-4918-a2c7-8a4e8d87ea3a · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Davis, Kristen Grauman, and Rog´erio Schmidt Feris
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c824db95-0761-4a49-8aef-c606a9170226 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Deep learning for video classification and captioning
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b0b61000-7204-46e3-9c89-70cbe5eb90f7 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Msr-vtt: A large video description dataset for bridging video and language
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b417a12e-c613-4623-b461-5dcfb2fc48c0 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MSR-VTT: A large video description dataset for bridging video and language
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 580d49b8-86fd-4237-b6e6-0aaf0f50a8b1 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0886b48f-09a5-4d1f-a63a-4601f2da034d · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding CLEVRER: CoLlision Events for Video REpresentation and Reasoning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86ed1b74-7c25-43e6-9888-5a20b69ec836 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Tenenbaum
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ac5110db-7965-434d-a3c8-be105974b135 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cbfdf1e-ed6a-4579-8036-f4c7b9d9ea87 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f67af649-7e95-4e51-9057-9c8392cd8a03 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Llama- adapter: Efficient fine-tuning of language models with zero- init attention
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f7d1270e-acba-4a7c-b2a9-061431d94cc9 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Please Carefully Think
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 831ab334-8259-405f-a2cb-5c78bfdc1331 · outbound
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work
Reference 200
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
No inbound Pith citation observations are available.