Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:38:20.743420Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 0 inbound Pith citation observations for arXiv:2507.10302.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:38:20.743420Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
92 of 92 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c201e185-eee5-49ee-a9b0-c2ca80d845f4 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43efa36a-bb39-4e18-8419-7712a080ed51 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2022) 1, 2
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 724bd4d6-c78a-45d0-a194-cfe9edf3479b · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f9a2c07-49b6-4862-a614-6beafbde6294 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICCV (2021) 2, 6, 9
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 607be67e-a3c6-44a2-8956-93aab90f7d4d · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2020) 2
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03173b9b-2712-4504-8da1-8ca05b4a22fe · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2020) 4
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7836d8d4-84b7-47c8-a079-162823de13a0 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1, 2
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f3e1cfc-98a5-41bd-9d90-15aea3596228 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 345d4ff9-0653-4faf-8052-39fb18424663 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66fe3c21-9130-4585-a951-609915ea4b3a · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2da34952-7e06-42c6-ad0f-7a60dfe52fc0 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86b558ce-f96b-4733-9f88-c59990bc0dd7 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs I see you, Batman!
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c4a722a-da2d-42d3-bc83-d5940e9691a8 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5498de92-18fa-4680-97ac-b93c70780507 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 3
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7ae5fcc-f264-4131-9fb5-268430b92ae1 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0194408-f5ec-4a50-9cf6-5439a4022bcf · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eef3aac-44ed-4ba3-bf11-518e7d03ccfa · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eefbd9ad-4aaf-413d-9aaa-49f560c5227e · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb31d280-a509-4892-8e52-0cb174b0c121 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41c84606-e81a-430b-9fc2-b1a174403ae8 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs PaLM-E: An Embodied Multimodal Language Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cd5c9eb-8562-4209-a371-08b49ddf1272 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs The Llama 3 Herd of Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2823f80-b137-44c5-b1ac-0516804f4955 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3776ed1-c023-4dcb-b04b-50b221e334e1 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2023) 6
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ce230bb-953d-4a1e-8900-a855f7c56296 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff4062fc-f51c-4b6a-8b42-aeb20bd31974 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Planting a SEED of Vision in Large Language Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02d3d1dd-4791-464b-b24d-bc252e7e3746 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICCV (2017) 6
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a15b683-b156-4d4c-9959-6d4ed0728984 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2022) 6
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 607dafda-84ce-4743-9e65-46991d82fbb3 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2019) 2
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e98d4342-9857-4619-b6b2-377dac9c3b79 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Mistral 7B
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c7f589a-3863-4465-a417-2f691152cc5e · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 6
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6cdc7be4-f0c6-42f2-8e3d-3274e708f59b · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs The Kinetics Human Action Video Dataset
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0959ef81-d36f-470e-82ac-a4902e94d8aa · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 1, 2
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 538bee68-24e6-4c16-ae51-3ecc023ed6df · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b348445d-e214-4a77-a551-d348e98fc3e9 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Otter: A Multi-Modal Model with In-Context Instruction Tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d59a25e4-0fa2-4725-9687-d6d4acfc89ff · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4520737c-1b66-4709-bf54-2658a49e8343 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2023) 1, 2, 3
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7542d54d-f7fc-4283-b69d-2ada832e5bca · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2022) 4, 6
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c86479a-1c60-4aa1-b402-216da35395b0 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VideoChat: Chat-Centric Video Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8289b59-9e87-4593-b8c3-81e9643c4161 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1, 2, 5, 6
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 442b2771-6305-4adf-85c8-671c77800522 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3243f840-cfe4-4a19-8f2d-3f79e8bfccb5 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 321d4b0d-5c05-495f-80a2-096a7362fa8d · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 5, 6
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32e494bc-9636-4bc8-9cc8-1b200e7e6b36 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 883d0dd4-0754-4599-8909-105cbb6e0d87 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2014) 2
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e629ad1a-c779-4479-8d24-7deba78bd179 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a44d768-9562-4cc9-afbf-b67f35e7967f · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2023) 1, 2, 3, 6
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f964ad-84e6-4297-9683-bfdb99e9498c · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 5, 6, 9
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8461f797-6b66-4aca-a0ad-f46468ac7f1f · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2020) 3
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d6b7801-4119-411d-8ec9-75f68c7a6da9 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2023) 2
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5bdbdf56-eabd-4041-96e7-58489e03b796 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e14157d-7c11-4d81-a775-cbeed9b6125d · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67b073d9-1d6a-4165-9ec4-b8d6bcc8e4ac · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2023) 6
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 386b79ff-461c-439c-8740-a77dfb464328 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 1
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49a829dd-f3dd-4985-bef6-b4571903db11 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 6
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee7759d0-0519-47ec-96a9-a2904642e3e4 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICCV (2015) 2
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43c0b533-4ac5-4a04-9d05-714a5d021326 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Streaming Long Video Understanding with Large Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7233cdfa-b97e-478c-92fb-9754b71359eb · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2021) 4
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7ffccc3-a817-4768-aebb-b72c3e2920ec · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a05af73-cc51-4fa8-a9fe-06b5697fa399 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2017) 3
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d2417f6-8f7b-4124-ad1b-50160aa41901 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbe2b09d-1871-496c-8ba4-195f40160540 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 3
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0e2a374-33f4-4108-a157-64655cfd577d · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs arXiv:2409.02889 (2024) 5
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9e902cd-cb82-47cf-82d4-603c9f294408 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 1, 2, 5, 6, 9
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 511b4426-8679-44f0-a05d-b86c191e34b7 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2021) 2, 6
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0de2c905-712d-4d08-b347-fd91b045fc66 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2021) 6 12
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7d431bf-cb92-4640-a746-f6f2056a90d8 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Slot-VLM: SlowFast Slots for Video-Language Modeling
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebaffa45-cddf-4b24-9b6d-ca4799c54f16 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2016) 2
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8da4eebb-4466-491b-b306-00581c9b7091 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34460155-d1c6-48db-85f6-23598d400956 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs RAL (2024) 1
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7bf3ec7f-6b4a-4085-a110-40881a5299fd · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5b3b342-ae0a-44a5-b0b4-153c0da017fd · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e0faf98-bfa6-4278-8748-473a2e7926dc · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1, 3
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81106527-a8cd-4c8f-a833-43e854334668 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2020) 6
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0beed7bf-d1b4-4ce8-8fe0-58119cbd0be3 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2024) 2
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0231cead-6e59-41bb-b365-ccd9cc2b50bf · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2016) 2
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc52b499-29e2-4fb4-a2ff-eff8f96974b8 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2023) 2
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3726dca5-5270-440f-946b-40c29d849afe · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c636545-9057-455b-86a3-c57df509d1d2 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf98a928-97fb-49f1-a8b2-0426a301e87f · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc289c9d-427d-4610-8baa-f295ed061a79 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e819ecc3-4df0-474b-b1df-8a031c12164b · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2025) 5, 6
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4bc26048-147c-4687-835c-a95a3e75083a · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 904869a4-664e-4648-b434-788a107d910f · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17b80f1f-639d-4ed0-8004-9db7e39d7fca · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cc34e76-338a-432b-b6d5-791c97873d5a · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 291812b2-9a35-4724-8333-e582821e6ec1 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2023) 2
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3aa8ffb8-1d66-486b-9788-233ae4cc58c3 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 1
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d69fd75a-9aa7-4bd6-928c-b428aaaf0e05 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ViLLa: Video Reasoning Segmentation with Large Language Model
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e973c4df-534a-4661-8449-cf606e65c9f5 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MLVU: Benchmarking Multi-task Long Video Understanding
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8ea2821-4834-40c5-8833-fc92a11bd334 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: AAAI (2018) 2
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4144ffb2-608f-478b-b583-7ffa57eae19f · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7f5f831-62fd-475b-bc0d-3474d9e72514 · outbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.