Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T07:08:35.946669Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 100 inbound Pith citation observations for arXiv:2406.16852.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T07:08:35.946669Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:57:40.232042Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
94 of 94 outbound references displayed
External citation measurements
2
pith, observed 2026-08-05T02:28:24.338817Z
Observation c366485e-fc67-4d54-95ac-b43e3ba6b31c · outbound
Long Context Transfer from Language to Vision Llm testneedleinahaystack
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 50743afc-919a-4f25-b93d-a8717ed4bf0c · outbound
Long Context Transfer from Language to Vision Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 659fe65a-0d19-44bc-80f1-ee50f65e057d · outbound
Long Context Transfer from Language to Vision Vqa: Visual question answering
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4623201c-56f0-44f2-86ae-8bb725a7a67c · outbound
Long Context Transfer from Language to Vision OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f2501790-375e-412d-8855-0935e1a217f7 · outbound
Long Context Transfer from Language to Vision Longalign: A recipe for long context alignment of large language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ecac27f0-511d-42f0-8085-7ce36732e94f · outbound
Long Context Transfer from Language to Vision ntkaware scaled rope allows llama models to have
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6d758d12-f987-4289-bbec-c07c07d1faff · outbound
Long Context Transfer from Language to Vision Ntk-aware scaled rope allows llama models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 905e55d7-98cd-4c23-9dc4-6e7238239d85 · outbound
Long Context Transfer from Language to Vision Language models are few-shot learners
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5ca33b95-7cad-422d-a50d-64fa184583d7 · outbound
Long Context Transfer from Language to Vision Matryoshka multimodal models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bf1bfdc5-d2bc-4d63-a79f-b38aa19b3a16 · outbound
Long Context Transfer from Language to Vision cerebras slimpajama-627b
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a489bee3-2c34-48ed-acc4-12e190345998 · outbound
Long Context Transfer from Language to Vision An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 220cbefe-a31b-48ed-af1c-e2c36a307709 · outbound
Long Context Transfer from Language to Vision Sharegpt4video: Improving video understanding and generation with better captions
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ed4463ef-97d7-4ee2-ad65-4b7cea21d963 · outbound
Long Context Transfer from Language to Vision Extending Context Window of Large Language Models via Positional Interpolation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c84878b0-bf24-47dc-b82a-86533c7295bf · outbound
Long Context Transfer from Language to Vision InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 93a0653b-d59c-4bce-a1fe-4d7fbac8707f · outbound
Long Context Transfer from Language to Vision Videollama 2: Advancing spatial- temporal modeling and audio understanding in video-llms
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c7426fe1-5d31-4fce-b4e6-9a468340e314 · outbound
Long Context Transfer from Language to Vision Generating long sequences with sparse transformers
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 56417ca4-2c47-41c7-9a90-51ecd373eaab · outbound
Long Context Transfer from Language to Vision Introducing command r+: A scalable llm built for business
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9d2e526e-1e09-453c-9c98-eebfe1bd5f4a · outbound
Long Context Transfer from Language to Vision Instructblip: Towards general-purpose vision-language models with instruction tuning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ee4847fd-1b8b-4220-a6be-7b7a231fbc9b · outbound
Long Context Transfer from Language to Vision Flashattention-2: Faster attention with better parallelism and work partitioning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4b2960b7-7bc7-4481-9f00-ce19b2ce278e · outbound
Long Context Transfer from Language to Vision LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bb0033e1-04b0-4d9c-a744-711677d53bb6 · outbound
Long Context Transfer from Language to Vision Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ccb39f39-a40e-484c-9d15-d677930e0011 · outbound
Long Context Transfer from Language to Vision Data engineering for scaling language models to 128k context
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 797d774f-cc90-4cb2-8c8c-b08ef8b7f0e1 · outbound
Long Context Transfer from Language to Vision Llmtest needleinahaystack
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 08ee7fe6-3659-4de2-8c9a-c843b8d993a9 · outbound
Long Context Transfer from Language to Vision Agqa: A benchmark for compositional spatio-temporal reasoning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b59f8ffe-9c0e-4f6f-8456-14319553b247 · outbound
Long Context Transfer from Language to Vision Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c59119fb-e4b1-46b1-bd9f-e2a18e02848a · outbound
Long Context Transfer from Language to Vision Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 875ff106-2979-44c8-b1aa-406718bb2297 · outbound
Long Context Transfer from Language to Vision Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 431282a5-8059-404a-9b69-e2a2f43a0b22 · outbound
Long Context Transfer from Language to Vision Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8ccc644b-33b7-4d06-8459-36c890c5d0b3 · outbound
Long Context Transfer from Language to Vision Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5871f125-0798-467b-bab6-ae427a9a2ca4 · outbound
Long Context Transfer from Language to Vision Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0877590c-f5dd-4a73-9700-8486f7cc4447 · outbound
Long Context Transfer from Language to Vision Chat-univi: Unified visual representation empowers large language models with image and video understanding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aa696503-b6b5-40f9-b092-e2abb5276edb · outbound
Long Context Transfer from Language to Vision A diagram is worth a dozen images
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ad2ea75d-fdb9-4749-8943-b4db2459357b · outbound
Long Context Transfer from Language to Vision Gonzalez, Hao Zhang, and Ion Stoica
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6f1e1d1a-ad11-4aa1-9ad8-d9e3780d8638 · outbound
Long Context Transfer from Language to Vision Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4148427a-4e00-4bf7-9dc7-4495c96d2938 · outbound
Long Context Transfer from Language to Vision Llava-next: What else influences visual instruction tuning beyond data?, May 2024
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9d162a63-2d51-4649-84b6-4574efb6e34d · outbound
Long Context Transfer from Language to Vision Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 31f68640-8092-47cc-919c-748e02322fe8 · outbound
Long Context Transfer from Language to Vision Otter: A multi-modal model with in-context instruction tuning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9ccd396b-2634-4d0b-92b0-0547650d9813 · outbound
Long Context Transfer from Language to Vision Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3e0bd944-f4ae-4368-b5de-b7ab0728c3cc · outbound
Long Context Transfer from Language to Vision Videochat: Chat-centric video understanding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1068bac3-20bc-41c1-90e9-06c95207df4d · outbound
Long Context Transfer from Language to Vision Sequence paral- lelism: Long sequence training from system perspective
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6187fefe-a1dd-4be2-9db1-77b4da39c36f · outbound
Long Context Transfer from Language to Vision Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dd08ca2c-2755-405e-bd68-230d44525070 · outbound
Long Context Transfer from Language to Vision Llama-vid: An image is worth 2 tokens in large language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 05ff61ff-ab17-4d8c-b015-49d4738e6899 · outbound
Long Context Transfer from Language to Vision Video-llava: Learning united visual representation by alignment before projection
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5313ed3a-bd0c-45ad-ae20-d32052be9d23 · outbound
Long Context Transfer from Language to Vision Vila: On pre-training for visual language models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bcf40407-7b27-431e-a3a0-41c56003b6cc · outbound
Long Context Transfer from Language to Vision World model on million-length video and language with ringattention
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6ba7ff8e-32a8-44f6-a7a6-327a4b115342 · outbound
Long Context Transfer from Language to Vision Ring attention with blockwise transformers for near-infinite context
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b1937ddf-b590-47ab-a549-e1ec801731e6 · outbound
Long Context Transfer from Language to Vision Improved baselines with visual instruction tuning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3cd987c5-616c-420a-8aed-1b523a984114 · outbound
Long Context Transfer from Language to Vision Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eacd7ac8-72f0-4f01-b779-7ca5b4c29200 · outbound
Long Context Transfer from Language to Vision Visual instruction tuning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation edd700d0-bac1-4967-8998-027020572c2c · outbound
Long Context Transfer from Language to Vision St-llm: Large language models are effective temporal learners
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d4c7d364-6ab5-4c3d-a24f-00d27f75c165 · outbound
Long Context Transfer from Language to Vision Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a079384e-8f2a-4e1e-95ed-6bebe30694ea · outbound
Long Context Transfer from Language to Vision Egoschema: A diagnostic benchmark for very long-form video language understanding
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 94a7c3d4-7e31-4ac8-a5cc-9bf6284d3177 · outbound
Long Context Transfer from Language to Vision Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9fc75294-2741-4978-a25d-8c0ab778eea9 · outbound
Long Context Transfer from Language to Vision DocVQA: A Dataset for VQA on Document Images
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a4947d69-546e-4380-bd39-e5541fb74f0a · outbound
Long Context Transfer from Language to Vision Mixtral 8x22b: Cheaper, better, faster, stronger
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c1bc88f2-a0fe-4a4b-b17b-341b0e9ff8bf · outbound
Long Context Transfer from Language to Vision Hello gpt-4o
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f429b08d-65b1-41f1-9c39-98f38a7d72e6 · outbound
Long Context Transfer from Language to Vision Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3a0800b8-10ef-4189-b47e-8c78998a67ec · outbound
Long Context Transfer from Language to Vision Yarn: Efficient context window extension of large language models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 362b40f3-e9d5-4e64-b3a3-0416a3148c13 · outbound
Long Context Transfer from Language to Vision Learning transferable visual models from natural language supervision
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6c4ee3c3-bd57-4ac5-a831-05d94d02d326 · outbound
Long Context Transfer from Language to Vision Zero: Memory optimiza- tions toward training trillion parameter models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 22c56be9-862e-450a-9f82-e8c2896aeafa · outbound
Long Context Transfer from Language to Vision Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eca5eac4-d395-4a51-b8da-559c783345c0 · outbound
Long Context Transfer from Language to Vision Code llama: Open foundation models for code
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0a8a4f38-3623-4c03-885a-5fb996eda2b5 · outbound
Long Context Transfer from Language to Vision Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1674f7d2-ad0d-4d10-b2e2-2f60d6fddc22 · outbound
Long Context Transfer from Language to Vision Milebench: Benchmarking mllms in long context
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 583f157c-35fc-4d17-bc05-3f8668164a16 · outbound
Long Context Transfer from Language to Vision Moviechat: From dense token to sparse memory for long video understanding
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b7848974-2c5c-4136-a03c-02b33deed058 · outbound
Long Context Transfer from Language to Vision Roformer: Enhanced transformer with rotary position embedding
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c029ae3e-6fdb-44fc-bb6b-fc9c8dcd80a8 · outbound
Long Context Transfer from Language to Vision Gemini: A family of highly capable multimodal models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bac81060-eebc-4633-b315-0f6d9ccc293b · outbound
Long Context Transfer from Language to Vision Palm 2 technical report
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 890e38de-801d-4718-939d-7bebca8c8dfd · outbound
Long Context Transfer from Language to Vision Introducing qwen-vl
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1d8b253b-215e-4eff-ad1b-67c920b7621e · outbound
Long Context Transfer from Language to Vision Qwen2 technical report
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cd07319c-223f-42b7-bcd9-21871d48c9dd · outbound
Long Context Transfer from Language to Vision Llama: Open and efficient foundation language models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 60c8d302-c1c5-4a5f-90a3-d4a61b5f0a8b · outbound
Long Context Transfer from Language to Vision Multimodal needle in a haystack: Benchmarking long- context capability of multimodal large language models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c89d548b-72b7-40db-a0ec-f0fee689e4f6 · outbound
Long Context Transfer from Language to Vision Lvbench: An extreme long video understanding benchmark
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7de1d54a-32a6-4c68-beba-338643ade4f5 · outbound
Long Context Transfer from Language to Vision Needle in a multimodal haystack
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 815cf966-3643-4aa2-b0c2-96a36dd569ef · outbound
Long Context Transfer from Language to Vision Star: A benchmark for situated reasoning in real-world videos
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d27cf5b0-d787-4d81-901f-26b4cdd0c1fc · outbound
Long Context Transfer from Language to Vision Grok-1.5 vision preview, apr 2024
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c9bf17de-a013-41a3-881a-48e082babbd4 · outbound
Long Context Transfer from Language to Vision Next-qa: Next phase of question- answering to explaining temporal actions
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 50eeabdb-7b33-4489-92a2-6b9a02b69970 · outbound
Long Context Transfer from Language to Vision Next-qa:next phase of question- answering to explaining temporal actions
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 46ec1724-4ac3-48eb-8de5-ec81857ad3aa · outbound
Long Context Transfer from Language to Vision Funqa: Towards surprising video comprehension
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d3f810b7-8414-48ae-893f-9cbc8e4c28e1 · outbound
Long Context Transfer from Language to Vision Effective long-context scaling of foundation models
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c2a5815c-bbda-47d8-a9c0-54cd7c814fc0 · outbound
Long Context Transfer from Language to Vision Video question answering via gradually refined attention over appearance and motion
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c44cc93b-e53a-4149-8024-8f5617e657aa · outbound
Long Context Transfer from Language to Vision Sutd-trafficqa: A question answering benchmark and an efficient network for video reasoning over traffic events
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2f87c9ed-ce08-4778-816f-5ac8eca9957e · outbound
Long Context Transfer from Language to Vision mplug-owl: Modularization empowers large language models with multimodality
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1217a881-608e-4e7d-b897-e1f6cfcf5548 · outbound
Long Context Transfer from Language to Vision CLEVRER: CoLlision Events for Video REpresentation and Reasoning
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 44614c79-4a71-45db-923e-2a8c9200fa7e · outbound
Long Context Transfer from Language to Vision Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d2a1a64a-f7d7-4198-8486-7ecd377754d9 · outbound
Long Context Transfer from Language to Vision Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1010d5e6-5683-4497-b6c1-516232eac650 · outbound
Long Context Transfer from Language to Vision Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 16d9716a-4914-4a9c-9aa2-e5951806d5ff · outbound
Long Context Transfer from Language to Vision Video-llama: An instruction-tuned audio-visual language model for video understanding
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6d7b7e90-727c-4797-8520-f6de88b6d6bb · outbound
Long Context Transfer from Language to Vision Direct preference optimization of video large multimodal models from language model reward
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f209fbee-f0be-4c60-b5a1-79fbef8c552d · outbound
Long Context Transfer from Language to Vision Llava-next: A strong zero-shot video understanding model, April 2024
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 67836fc9-a424-498c-b75b-ecf15dfc5b77 · outbound
Long Context Transfer from Language to Vision Mlvu: A comprehensive benchmark for multi-task long video understanding
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 01160dc2-3cda-4ca2-8568-cf65190e437e · outbound
Long Context Transfer from Language to Vision Streaming dense video captioning
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bd86627d-912b-4fed-94ca-7b72dffa1ecf · outbound
Long Context Transfer from Language to Vision Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4ff52b4b-5133-43f6-8835-03c7cf80f707 · outbound
Long Context Transfer from Language to Vision Ring flash attention
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1251f619-4ad9-4d84-8eba-de96e809bfcb · inbound
MLVU: Benchmarking Multi-task Long Video Understanding Long Context Transfer from Language to Vision
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3929cb71-7217-45bb-9bbc-5cb5361efda5 · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Long Context Transfer from Language to Vision
Reference 175
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2a30ea10-0ddd-4fa5-8371-e8240ea9ff57 · inbound
LLaVA-OneVision: Easy Visual Task Transfer Long Context Transfer from Language to Vision
Reference 163
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1194b3af-a1bb-45c0-a5a6-b9f2c4766120 · inbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Long Context Transfer from Language to Vision
Reference 156
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 69c4ae92-980c-4438-bcce-fa92c87cbe02 · inbound
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding Long Context Transfer from Language to Vision
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e9fb2502-0eb6-4e15-a7ff-d8cb773cf645 · inbound
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces Long Context Transfer from Language to Vision
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e272bae0-47ed-44e8-9638-54b01db80967 · inbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Long Context Transfer from Language to Vision
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1d0859dd-9a8e-4160-865c-1f35042ff9c3 · inbound
Online Video Understanding: OVBench and VideoChat-Online Long Context Transfer from Language to Vision
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa6abe33-b726-44dc-a3f6-008cf740de43 · inbound
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM Long Context Transfer from Language to Vision
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0790486-e797-4a9b-a1d8-d0c314216d93 · inbound
HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding Long Context Transfer from Language to Vision
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85cd279f-e32d-49ea-bb59-543be1c4b444 · inbound
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction Long Context Transfer from Language to Vision
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5e4720b0-756f-4074-a3ec-614e11c3be19 · inbound
LongViTU: Instruction Tuning for Long-Form Video Understanding Long Context Transfer from Language to Vision
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdfd847a-ef76-48f1-a66e-f25636e27b09 · inbound
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding Long Context Transfer from Language to Vision
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c62dc1ae-7bce-4b19-8ef8-a581782e4001 · inbound
NExtLong: Toward Effective Long-Context Training without Long Documents Long Context Transfer from Language to Vision
Reference 113
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 987e8f26-1a31-4758-85dc-5e9458ed81d4 · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Long Context Transfer from Language to Vision
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f5040c62-020c-4353-bab4-32171b018be4 · inbound
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge Long Context Transfer from Language to Vision
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c28d9e8e-ea83-486e-a407-c8508b62acbc · inbound
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Long Context Transfer from Language to Vision
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cfb261d3-c08a-4034-8835-35aad62cee14 · inbound
Temporal Preference Optimization for Long-Form Video Understanding Long Context Transfer from Language to Vision
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccba3e18-c063-432f-a6f1-204bd7e412ce · inbound
$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation Long Context Transfer from Language to Vision
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 078991f5-1663-4b24-ac01-f52c2f54a09a · inbound
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos Long Context Transfer from Language to Vision
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e79ee3c6-d068-4c32-9b59-2ac9a81cf4ef · inbound
HD-EPIC: A Highly-Detailed Egocentric Video Dataset Long Context Transfer from Language to Vision
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d38a661a-07d0-4337-94fe-949196180adb · inbound
CoS: Chain-of-Shot Prompting for Long Video Understanding Long Context Transfer from Language to Vision
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5facfe1c-fbc9-40fe-938b-35110551efa7 · inbound
LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation Long Context Transfer from Language to Vision
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e99c838d-a56e-425e-8168-ddef02a2589a · inbound
Video-R1: Reinforcing Video Reasoning in MLLMs Long Context Transfer from Language to Vision
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ada76126-f8e5-4621-8c0b-fd110515a6dd · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Long Context Transfer from Language to Vision
Reference 146
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fbb17e75-ffea-475a-947c-10ab364bd87b · inbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Long Context Transfer from Language to Vision
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 686d382c-f69d-4325-9780-7c48d5248f4d · inbound
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Long Context Transfer from Language to Vision
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 896f522a-1fd3-4c1f-9435-e1462eb735e9 · inbound
Clapper: Compact Learning and Video Representation in VLMs Long Context Transfer from Language to Vision
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c45523ae-bfe6-4356-b6b4-3f7ce6e92d7d · inbound
Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Long Context Transfer from Language to Vision
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 094c0bcc-ac62-41ab-9ccc-b719cc5beb56 · inbound
Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought Long Context Transfer from Language to Vision
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2108de4d-6b3d-48a7-8140-aff954899ef2 · inbound
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Long Context Transfer from Language to Vision
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af4007e5-3f71-456d-80e7-b3c04993e6b8 · inbound
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Long Context Transfer from Language to Vision
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation df459411-a9ef-4025-ba51-db9a32e0e78d · inbound
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Long Context Transfer from Language to Vision
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d0f16f33-24b5-4af9-ad4e-4b8e07355c6d · inbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Long Context Transfer from Language to Vision
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 708677d5-b6a1-4e7b-88c6-5af249e4da12 · inbound
Reinforcing Video Reasoning with Focused Thinking Long Context Transfer from Language to Vision
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0efa8bca-4d4f-4abb-a51f-dcad74ffb272 · inbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Long Context Transfer from Language to Vision
Reference 106
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31892af6-6997-4b2c-9b72-0e9cc0709858 · inbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Long Context Transfer from Language to Vision
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ac67125-e712-4a7e-ac56-d9575d4c0163 · inbound
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding Long Context Transfer from Language to Vision
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc55fd6e-4219-447b-b325-f6fe6f62535c · inbound
Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Long Context Transfer from Language to Vision
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0a8dfb3-6cc0-40ee-a393-04b366ffd805 · inbound
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence Long Context Transfer from Language to Vision
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7090f9e5-e826-48d2-a982-117d9ec7574c · inbound
Vid-SME: Membership Inference Attacks against Large Video Understanding Models Long Context Transfer from Language to Vision
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b10c4db0-dc68-4031-b0c2-9d98e407ecf4 · inbound
TextVidBench: A Benchmark for Long Video Scene Text Understanding Long Context Transfer from Language to Vision
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d56b795a-64cc-4291-86fd-07789b11d2f0 · inbound
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs Long Context Transfer from Language to Vision
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2a45ff1-3772-427d-a662-2d50614b02a3 · inbound
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? Long Context Transfer from Language to Vision
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d35a5bff-efa6-4b40-aa4c-bba752662cc4 · inbound
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos Long Context Transfer from Language to Vision
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 372c0c6c-ddb6-46d8-8879-8a2903415c3f · inbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Long Context Transfer from Language to Vision
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf0663c6-07a0-424d-a475-16f65e101b6b · inbound
CyberV: Cybernetics for Test-time Scaling in Video Understanding Long Context Transfer from Language to Vision
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48c77cfe-290f-4270-adb6-2b1adcf3f0d6 · inbound
Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences? Long Context Transfer from Language to Vision
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 780e01ce-6412-459e-9e7b-5403e03752ce · inbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Long Context Transfer from Language to Vision
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26b7f326-3691-4767-a429-eddfeaa91aef · inbound
Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Long Context Transfer from Language to Vision
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d807b7f7-e6c9-475d-93ee-a9cd1cf256d3 · inbound
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning Long Context Transfer from Language to Vision
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81ae1300-a353-4646-9fc7-5186e7de2a7a · inbound
Show-o2: Improved Native Unified Multimodal Models Long Context Transfer from Language to Vision
Reference 143
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0b8f8ce4-9400-4786-81bd-7c484317d5cc · inbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Long Context Transfer from Language to Vision
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35bc73c8-054a-4132-919a-d3a9d7e7a3e6 · inbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Long Context Transfer from Language to Vision
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38fdb9fe-d313-4570-8f2b-565d97b42853 · inbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Long Context Transfer from Language to Vision
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 122a666a-3a86-4f3c-a925-119f5f0a1358 · inbound
Task-Aware KV Compression For Cost-Effective Long Video Understanding Long Context Transfer from Language to Vision
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef8de3d9-0280-4974-bb2f-47853121def6 · inbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Long Context Transfer from Language to Vision
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac77033f-768b-4197-b7c0-04d0923444c2 · inbound
MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Long Context Transfer from Language to Vision
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fafdc095-d006-47ba-8474-889537b0e4ff · inbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Long Context Transfer from Language to Vision
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 079be2da-dd1c-462c-b917-a51a2c231eee · inbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Long Context Transfer from Language to Vision
Reference 112
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 602199ed-14c7-4fbf-9815-939352dc8fa4 · inbound
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Long Context Transfer from Language to Vision
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6118feca-54de-4a13-896f-92b60e129dd8 · inbound
RadEyeVideo: Enhancing general-domain Large Vision Language Model for chest X-ray analysis with video representations of eye gaze Long Context Transfer from Language to Vision
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 009ac0f9-9498-40c2-b9cd-b9a93ba48836 · inbound
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models Long Context Transfer from Language to Vision
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe22ba31-6ef0-422e-a89e-4b5188af112e · inbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Long Context Transfer from Language to Vision
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b5e3be6-f9fb-455b-8214-d751b065597f · inbound
LAVA: Language Driven Scalable and Versatile Traffic Video Analytics Long Context Transfer from Language to Vision
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3db7dd76-d71f-41af-b576-3b93ebe9c2ee · inbound
Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection Long Context Transfer from Language to Vision
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation defef782-e812-4bf7-8dcb-9fda3c5ddc59 · inbound
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding Long Context Transfer from Language to Vision
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 217ed845-1e01-4a96-a6ec-079d028f42d4 · inbound
Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment Long Context Transfer from Language to Vision
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a75667df-189b-40d5-a2e7-ab369304ba27 · inbound
DATE: Dynamic Absolute Time Enhancement for Long Video Understanding Long Context Transfer from Language to Vision
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7648395d-473f-4b5d-a0fc-713faf085d21 · inbound
Video Reasoning without Training Long Context Transfer from Language to Vision
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 738943dd-7f6f-4f19-83b6-12b008fa9194 · inbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Long Context Transfer from Language to Vision
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4504deb-bbe8-4d43-bf54-8eddaaf3a9c8 · inbound
Cambrian-S: Towards Spatial Supersensing in Video Long Context Transfer from Language to Vision
Reference 160
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b6d85b97-3627-4be6-b44e-c209e842c7f1 · inbound
VIDEOP2R: Video Understanding from Perception to Reasoning Long Context Transfer from Language to Vision
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5d0bb3f1-e85d-4ecf-bb20-9a521d582944 · inbound
REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding Long Context Transfer from Language to Vision
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7b972113-b370-45c8-9b32-5690607e6a70 · inbound
Vision-Language Memory for Spatial Reasoning Long Context Transfer from Language to Vision
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82e609f1-d849-4451-89ca-dde64e8fbb9e · inbound
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Long Context Transfer from Language to Vision
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f964a623-93b3-404c-b769-131bde129cf3 · inbound
Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval Long Context Transfer from Language to Vision
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ed169029-a8de-464c-8137-bea9c7cca2a2 · inbound
SpatialMosaic: A Multiview VLM Dataset for Partial Visibility Long Context Transfer from Language to Vision
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ff464fde-2c93-413f-85b0-02642e352f9f · inbound
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding Long Context Transfer from Language to Vision
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 334625a9-0400-42c3-8049-4db0e5b16ca5 · inbound
MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources Long Context Transfer from Language to Vision
Reference 131
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4780408a-6c60-4b1f-b636-5a1de8a8f82d · inbound
ReMoT: Reinforcement Learning with Motion Contrast Triplets Long Context Transfer from Language to Vision
Reference 123
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1862eb9f-6e6d-4eac-9e5b-43713013f906 · inbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Long Context Transfer from Language to Vision
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 79630530-51c6-4ae5-b6c6-df80086870f8 · inbound
From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation Long Context Transfer from Language to Vision
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7d8a661-922b-44cd-81cc-fdea29c5d0af · inbound
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding Long Context Transfer from Language to Vision
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b78802e-815f-412f-89dc-f2e33e3ba61a · inbound
Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark Long Context Transfer from Language to Vision
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bf500279-c1a0-4dad-b808-c81cae10a01d · inbound
STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering Long Context Transfer from Language to Vision
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4fb48766-c116-4ff5-ba90-e5f6da5eae34 · inbound
Let Geometry GUIDE: Layer-wise Unrolling of Geometric Priors in Multimodal LLMs Long Context Transfer from Language to Vision
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fe8f2b3f-4c31-4421-8f8e-7b6b1f9de0ee · inbound
HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models Long Context Transfer from Language to Vision
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c7f6e111-2efb-435c-b1fa-f3126e2d2c39 · inbound
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding Long Context Transfer from Language to Vision
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0d36e27f-b624-4232-8a0a-6456a6de7a6a · inbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding Long Context Transfer from Language to Vision
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a2d7224b-1660-4b29-a8b2-0640f6d3c36e · inbound
Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging Long Context Transfer from Language to Vision
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dde20986-fddd-4353-96d0-49af6583fab0 · inbound
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Long Context Transfer from Language to Vision
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation edf47762-9789-499e-95b4-89092e306270 · inbound
SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing Long Context Transfer from Language to Vision
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0d825d63-2330-46e5-beb2-60a149bad419 · inbound
OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning Long Context Transfer from Language to Vision
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ed264919-1890-4210-a810-194c5ecf7f45 · inbound
EgoSelf: From Memory to Personalized Egocentric Assistant Long Context Transfer from Language to Vision
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2ef0ece3-f894-4748-8c2a-b7124869c768 · inbound
Video-ToC: Video Tree-of-Cue Reasoning Long Context Transfer from Language to Vision
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fa674b76-523d-4ce2-96ac-602a24c065ff · inbound
HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration Long Context Transfer from Language to Vision
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dfb2920d-67fc-491a-b61a-b28eb6e5d18b · inbound
CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding Long Context Transfer from Language to Vision
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5b998405-1ec4-4dd2-b88e-ad2fa3c56b75 · inbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Long Context Transfer from Language to Vision
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 441a83b3-bcfa-43fe-9b7c-3c71bebbe125 · inbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Long Context Transfer from Language to Vision
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.