Pith. sign in

Paper Citation Record · LEDGER

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs

As of 7 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 0 inbound Pith citation observations for arXiv:2507.10302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10302 v1

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:38:20.743420Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

92 of 92 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c201e185-eee5-49ee-a9b0-c2ca80d845f4 · outbound

This paper cites GPT-4 Technical Report.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:13.744309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:13.744309Z digest=sha256:596905278a5b05697e69887cf14ec3ba9abd4aa8df5c1609d70cd4603b2840ed

Observation 43efa36a-bb39-4e18-8419-7712a080ed51 · outbound

This paper cites In: NeurIPS (2022) 1, 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2022) 1, 2

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:13.800454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:13.800454Z digest=sha256:59f6024eb2264f534379afe65645442df05a8e1b1d15e6a6688feafd62c4486e

Observation 724bd4d6-c78a-45d0-a194-cfe9edf3479b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:13.887420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:13.887420Z digest=sha256:9202a6c313bec97a50cc68f6b037efcddbc3ce607a4c2db0478857b1124a0cdc

Observation 9f9a2c07-49b6-4862-a614-6beafbde6294 · outbound

This paper cites In: ICCV (2021) 2, 6, 9.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICCV (2021) 2, 6, 9

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:13.993701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:13.993701Z digest=sha256:ff600d19f199cab3c9d141b983884dcb6999bafc49a6ab19075e78aab69d9627

Observation 607be67e-a3c6-44a2-8956-93aab90f7d4d · outbound

This paper cites In: NeurIPS (2020) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2020) 2

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.035353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.035353Z digest=sha256:b361b3989fc67d1302c5f65bfd417ed86047d04d5318a85c37278a95bd74e2df

Observation 03173b9b-2712-4504-8da1-8ca05b4a22fe · outbound

This paper cites In: ECCV (2020) 4.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2020) 4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.105134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.105134Z digest=sha256:f9d924678fb27ac21db332d48f0afe6c1b01140281322b1b5be1a250b7bc6335

Observation 7836d8d4-84b7-47c8-a079-162823de13a0 · outbound

This paper cites In: CVPR (2024) 1, 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1, 2

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.175728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.175728Z digest=sha256:b399b81b86f421685f43ae9bd23629a019f5f2b7038b4d82137a93d1fc083fec

Observation 4f3e1cfc-98a5-41bd-9d90-15aea3596228 · outbound

This paper cites Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.260667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.260667Z digest=sha256:b24ea9e0cd98d300078b0e95fac9a6315deb56d2386309a9ecf05183217bdcf4

Observation 345d4ff9-0653-4faf-8052-39fb18424663 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.332299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.332299Z digest=sha256:cd4e7c00304822b2da8c4950992f96901edffb4d8629dcff936ba5b599c1dcb8

Observation 66fe3c21-9130-4585-a951-609915ea4b3a · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.390971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.390971Z digest=sha256:8eddfbbf0875f3413ba3b9ff2d5802dad8b6f39be209f3ab83597ab1c5ae8d36

Observation 2da34952-7e06-42c6-ad0f-7a60dfe52fc0 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.432529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.432529Z digest=sha256:6ac172d6acf364069558067ea9cbb9a810638363c1859e5b075502be638a0a78

Observation 86b558ce-f96b-4733-9f88-c59990bc0dd7 · outbound

This paper cites I see you, Batman!.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs I see you, Batman!

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.513288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.513288Z digest=sha256:4a7b597900c4e3779566b73e4ab09508c77afdf8835c9596f6acecac94016219

Observation 5c4a722a-da2d-42d3-bc83-d5940e9691a8 · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.586386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.586386Z digest=sha256:9b66c22b7889564a0d2e5c2e92e8650fa0af6b8b7a8842a9401ba80c46042b8e

Observation 5498de92-18fa-4680-97ac-b93c70780507 · outbound

This paper cites In: CVPR (2024) 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 3

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.671358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.671358Z digest=sha256:aef3b18b1609e5e14d953eb06b649d3f9d10f7016cee509ef5cb789885e13603

Observation b7ae5fcc-f264-4131-9fb5-268430b92ae1 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.763233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.763233Z digest=sha256:073a4f2607a7fbf2d0235984ac4abd81a7082b9a6ce792cb4dd870d775aa4c9d

Observation f0194408-f5ec-4a50-9cf6-5439a4022bcf · outbound

This paper cites an unresolved cited work.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.830921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.830921Z digest=sha256:69d32d3ed88e7bdce8cb44a3e3e8da80891c305f414ab64a87e418d235e1ea31

Observation 1eef3aac-44ed-4ba3-bf11-518e7d03ccfa · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.924271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.924271Z digest=sha256:c0142563e8ad40508f6500cfbc41c6b0226ede520612795e9eada698077f83a1

Observation eefbd9ad-4aaf-413d-9aaa-49f560c5227e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.022435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.022435Z digest=sha256:4aa9ef7f3aaf83debf41ebaa059d4b1622c71fca94979d1819ddd6bbcefc658b

Observation cb31d280-a509-4892-8e52-0cb174b0c121 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.079259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.079259Z digest=sha256:972268958d73a416f4da94ca3169ccf6bd9c1010817dc229d409852cb784d81c

Observation 41c84606-e81a-430b-9fc2-b1a174403ae8 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs PaLM-E: An Embodied Multimodal Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.184382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.184382Z digest=sha256:0cb61db40a32fdca9b2d0890f40f1762b8e5ad0fdbd7fde4e04cabebbf1f7d3a

Observation 8cd5c9eb-8562-4209-a371-08b49ddf1272 · outbound

This paper cites The Llama 3 Herd of Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.276356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.276356Z digest=sha256:a92dda8363b0d8666b96149fc65e3388a485a3b8f6d6a1b9d15f458d0c0809f1

Observation b2823f80-b137-44c5-b1ac-0516804f4955 · outbound

This paper cites an unresolved cited work.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.378346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.378346Z digest=sha256:6eba113fe422d4dc05d579c9ec333638530aa44093c1669428ef73a8e2dc6db5

Observation e3776ed1-c023-4dcb-b04b-50b221e334e1 · outbound

This paper cites In: CVPR (2023) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2023) 6

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.426339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.426339Z digest=sha256:673504a56c8c95a7a582acddf748586cfacc0deff4cf2c223b453296fa3efb62

Observation 0ce230bb-953d-4a1e-8900-a855f7c56296 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.473346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.473346Z digest=sha256:2973faa51642d1fecd54ab26d7643c87bbab2a197608ce786e3379371c6c9c19

Observation ff4062fc-f51c-4b6a-8b42-aeb20bd31974 · outbound

This paper cites Planting a SEED of Vision in Large Language Model.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Planting a SEED of Vision in Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.540771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.540771Z digest=sha256:2adf5296dd22d2a79e15485f0e18ca5b37273b9a10b3e81f8b63f379d03dfcc9

Observation 02d3d1dd-4791-464b-b24d-bc252e7e3746 · outbound

This paper cites In: ICCV (2017) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICCV (2017) 6

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.614673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.614673Z digest=sha256:9415cf9b970732d8094007d5e26733378e39837a88833972956c8500ad43d5a8

Observation 1a15b683-b156-4d4c-9959-6d4ed0728984 · outbound

This paper cites In: ICLR (2022) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2022) 6

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.687789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.687789Z digest=sha256:b09864e78dd9b91438899b77327a936977d8bd4117e642ced447a3aeb566deff

Observation 607dafda-84ce-4743-9e65-46991d82fbb3 · outbound

This paper cites In: CVPR (2019) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2019) 2

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:31.066212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:15.781525Z digest=sha256:21fdab072d650c19ed135a6b9f8819f07a7260180a25af2a3485af2b16869995

Observation e98d4342-9857-4619-b6b2-377dac9c3b79 · outbound

This paper cites Mistral 7B.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Mistral 7B

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.849245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.849245Z digest=sha256:cc06b0edfbdff1ae48d6bfcbb4a3d22dfa752691ed0352fa5e6fcd85d44b6309

Observation 2c7f589a-3863-4465-a417-2f691152cc5e · outbound

This paper cites In: CVPR (2024) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 6

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:30.848354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:15.963244Z digest=sha256:4a8f4cbe495761746a504efc2461e3a0409856e1a01b9e718b49217bfedc9b4d

Observation 6cdc7be4-f0c6-42f2-8e3d-3274e708f59b · outbound

This paper cites The Kinetics Human Action Video Dataset.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs The Kinetics Human Action Video Dataset

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.022354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.022354Z digest=sha256:227926c6765c743578ddcc473e566e9f3dac8c5884650c39db89cb2788a429bc

Observation 0959ef81-d36f-470e-82ac-a4902e94d8aa · outbound

This paper cites In: ECCV (2024) 1, 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 1, 2

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:30.531168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.087756Z digest=sha256:fc33264cc370145791e08e5685a0a9808d4d30b4e131687f026fd0c84b286a9f

Observation 538bee68-24e6-4c16-ae51-3ecc023ed6df · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.116489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.116489Z digest=sha256:c4251ee52c87d9ea08ef51b000fe6d04a4a58487194bc1f12ffe489c352331a4

Observation b348445d-e214-4a77-a551-d348e98fc3e9 · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.192030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.192030Z digest=sha256:f2c509a7dc8f8c6a9a2695e9f188e9544dc6815ec890c2c5ee7df9d62a289764

Observation d59a25e4-0fa2-4725-9687-d6d4acfc89ff · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.241699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.241699Z digest=sha256:fa6c194dc76baaa1f082f94e81e7ea19964f0059c2c5cc15c6212ba73d603523

Observation 4520737c-1b66-4709-bf54-2658a49e8343 · outbound

This paper cites In: ICML (2023) 1, 2, 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2023) 1, 2, 3

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:30.215814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.297011Z digest=sha256:5d0d65cbeb7dc3b5ec1c9623f5f344a4ddbf6b134b73fc71f7db2d36868d8c3e

Observation 7542d54d-f7fc-4283-b69d-2ada832e5bca · outbound

This paper cites In: ICML (2022) 4, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2022) 4, 6

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.902598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.371511Z digest=sha256:295dc14cc8a2e1fd720c0c42775686630d56956bec88291efb3cdd61956340f7

Observation 7c86479a-1c60-4aa1-b402-216da35395b0 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VideoChat: Chat-Centric Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.430652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.430652Z digest=sha256:a2ad8a31633ec799ff2806ed56b383255ec47d4d149b2f7739ea767cfdbbe70a

Observation d8289b59-9e87-4593-b8c3-81e9643c4161 · outbound

This paper cites In: CVPR (2024) 1, 2, 5, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1, 2, 5, 6

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.657002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.489156Z digest=sha256:aeb17d4357926104060e3a3f27200f7a4737e240cd07fd46f71f3a44d7fbba1d

Observation 442b2771-6305-4adf-85c8-671c77800522 · outbound

This paper cites TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T17:38:21.336547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.555603Z digest=sha256:ecc9a4068790b357704bfd1667623f4f02201673ad48a7be1537395334a16113

Observation 3243f840-cfe4-4a19-8f2d-3f79e8bfccb5 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.589632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.589632Z digest=sha256:ea95fb6fe080dcf812606ecdf07053613a06bf961460a76bd1acde580d3968dd

Observation 321d4b0d-5c05-495f-80a2-096a7362fa8d · outbound

This paper cites In: ECCV (2024) 5, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 5, 6

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.498383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.624560Z digest=sha256:55315cc6f5df94d78f2f36241ce4f6ef6c233001f5013f2c972859def1709a99

Observation 32e494bc-9636-4bc8-9cc8-1b200e7e6b36 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.675424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.675424Z digest=sha256:2c2520e6e6e061d5bc226983acc5081efa4eacbdb8c7c7742fdd9e3c745b1929

Observation 883d0dd4-0754-4599-8909-105cbb6e0d87 · outbound

This paper cites In: ECCV (2014) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2014) 2

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.304790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.723028Z digest=sha256:20222ad630671f7d09210a63467012d627e67ec2fd8efa0bf783914da03966a9

Observation e629ad1a-c779-4479-8d24-7deba78bd179 · outbound

This paper cites In: CVPR (2024) 1.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.062686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.765941Z digest=sha256:432a87aa14b1334228473eb7b9c4a09adc6498ef2f30e61739608cb81fda1d62

Observation 4a44d768-9562-4cc9-afbf-b67f35e7967f · outbound

This paper cites In: NeurIPS (2023) 1, 2, 3, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2023) 1, 2, 3, 6

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.846803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.846803Z digest=sha256:ec10079c3d5a420fd8a2cd38c370028249e5b351425e720c1d2d572a7b6809df

Observation 81f964ad-84e6-4297-9683-bfdb99e9498c · outbound

This paper cites In: ECCV (2024) 5, 6, 9.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 5, 6, 9

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:28.764386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.889577Z digest=sha256:561125c820ac7058c06ce5e1ef94e24efddf412bbdf16329d6b7b357d4b06bca

Observation 8461f797-6b66-4aca-a0ad-f46468ac7f1f · outbound

This paper cites In: NeurIPS (2020) 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2020) 3

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:28.539521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.916656Z digest=sha256:b8183d04fd02481ff3c77eb95c64ff650657657cce7170052b87684165d6301e

Observation 7d6b7801-4119-411d-8ec9-75f68c7a6da9 · outbound

This paper cites In: ICML (2023) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2023) 2

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:28.318729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.966244Z digest=sha256:ca7c7b733d6473af5925aa244c868cf53ab6b65c2ccd72f9d4806d3d897e074e

Observation 5bdbdf56-eabd-4041-96e7-58489e03b796 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.002664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.002664Z digest=sha256:942c3924882523369d358483b403beca6a67da1e145b4b250bbdc8662158d893

Observation 8e14157d-7c11-4d81-a775-cbeed9b6125d · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.072193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.072193Z digest=sha256:3e9fe6a6f70da2600b716c06ded644bcaa05968bd426f408b8968c40e9403adb

Observation 67b073d9-1d6a-4165-9ec4-b8d6bcc8e4ac · outbound

This paper cites In: NeurIPS (2023) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2023) 6

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:28.175772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:17.130992Z digest=sha256:a5b3e150f326258b0d877aa8f66d74276e8170d4cb079c268f9aba03d4a96768

Observation 386b79ff-461c-439c-8740-a77dfb464328 · outbound

This paper cites In: NeurIPS (2024) 1.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 1

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.903975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:17.161209Z digest=sha256:ce8be8fe3bdd3291d321872949d5fddbf35a91450a21d8fe220b6af6202454f8

Observation 49a829dd-f3dd-4985-bef6-b4571903db11 · outbound

This paper cites In: NeurIPS (2024) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 6

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.677008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:17.208737Z digest=sha256:f605b969529e9f3073b1364197b35a75cf4a08d684479c1c7e9c97b1ff197f90

Observation ee7759d0-0519-47ec-96a9-a2904642e3e4 · outbound

This paper cites In: ICCV (2015) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICCV (2015) 2

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.487305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:17.278269Z digest=sha256:8e80a5ce2f01f8e41162e8ed8905d702bc564e4e70edcaf46135ab67e97d377e

Observation 43c0b533-4ac5-4a04-9d05-714a5d021326 · outbound

This paper cites Streaming Long Video Understanding with Large Language Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Streaming Long Video Understanding with Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.320965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.320965Z digest=sha256:4ea86a1995a9ad1564c062623d2eaaccab0b60f5db7a77be8ee58d8584a4b2c4

Observation 7233cdfa-b97e-478c-92fb-9754b71359eb · outbound

This paper cites In: ICML (2021) 4.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2021) 4

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.263986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:17.439568Z digest=sha256:cd9a1fc46852388fc5b79f15a286b578db0a138964f46d25542b696cb635c97c

Observation b7ffccc3-a817-4768-aebb-b72c3e2920ec · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.535347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.535347Z digest=sha256:7a2633309df733f0691eef73605f805836e9e9d87cdcabfa5f9419bc66b541f2

Observation 6a05af73-cc51-4fa8-a9fe-06b5697fa399 · outbound

This paper cites In: NeurIPS (2017) 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2017) 3

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.011795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:17.677509Z digest=sha256:ad9c475335426c89e27cf8de3bf625c16457948f8848c4c0d75168dfec25240b

Observation 8d2417f6-8f7b-4124-ad1b-50160aa41901 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.726972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.726972Z digest=sha256:b1a6b52e41d984c5be4a76a59ada695813a8cfc507c925ed0fd69c0f1cba60d1

Observation fbe2b09d-1871-496c-8ba4-195f40160540 · outbound

This paper cites In: NeurIPS (2024) 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 3

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:26.719355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:17.788180Z digest=sha256:eca9becc63f34e1fbe26ce894183132ec5b68eaea9c051ad37107d1d5ca20386

Observation d0e2a374-33f4-4108-a157-64655cfd577d · outbound

This paper cites arXiv:2409.02889 (2024) 5.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs arXiv:2409.02889 (2024) 5

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.876527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.876527Z digest=sha256:ebaadaf5caeddf6bde86a96a4b2b67713be559a679c916cab3b2cde12e4ab8f4

Observation a9e902cd-cb82-47cf-82d4-603c9f294408 · outbound

This paper cites In: ECCV (2024) 1, 2, 5, 6, 9.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 1, 2, 5, 6, 9

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:26.539128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:17.935120Z digest=sha256:21772137e6f8e70686987f6e042ee5847c78c1d5469e10c285d8b03572ffb252

Observation 511b4426-8679-44f0-a05d-b86c191e34b7 · outbound

This paper cites In: NeurIPS (2021) 2, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2021) 2, 6

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:26.310005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:18.007029Z digest=sha256:8de112d946052017f9dcca1cc2fa884e70f5c7fb51e14f387ad91fa3bc70a4ec

Observation 0de2c905-712d-4d08-b347-fd91b045fc66 · outbound

This paper cites In: CVPR (2021) 6 12.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2021) 6 12

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:25.983174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:18.031278Z digest=sha256:847dcd8bcba9da7b09285659a2b61d8216a04160f81f5feba3e676421dbdd16e

Observation f7d431bf-cb92-4640-a746-f6f2056a90d8 · outbound

This paper cites Slot-VLM: SlowFast Slots for Video-Language Modeling.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Slot-VLM: SlowFast Slots for Video-Language Modeling

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.120829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.120829Z digest=sha256:3b8e8e2b9f0d842bc1c6325e99b158c70b093f26d0a432ceb18bec4f16bd531f

Observation ebaffa45-cddf-4b24-9b6d-ca4799c54f16 · outbound

This paper cites In: CVPR (2016) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2016) 2

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:25.628447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:18.244765Z digest=sha256:6f7745018bcecf86e87be2ca2a548b6f5027fdb87c408f877a709983b89d0c15

Observation 8da4eebb-4466-491b-b306-00581c9b7091 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.355450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.355450Z digest=sha256:77e90a4302d8826a6950f489cab454814769d8c05ef948dd00380afe4ebfac38

Observation 34460155-d1c6-48db-85f6-23598d400956 · outbound

This paper cites RAL (2024) 1.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs RAL (2024) 1

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:25.341643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:18.449200Z digest=sha256:f4fc799c23b17d29c58a9a59e4c8dd9d9a7f7693f671e7c5c61ab0c1e706e460

Observation 7bf3ec7f-6b4a-4085-a110-40881a5299fd · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.564455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.564455Z digest=sha256:401c3779cb6b4b652b2bfe6005cbbca685f2ebe72fa712efa29bd30cd482e605

Observation e5b3b342-ae0a-44a5-b0b4-153c0da017fd · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.659223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.659223Z digest=sha256:205d498e169cd1c0b12336ef266e387c325a380225b8af350c9c767a9fc4c8ad

Observation 8e0faf98-bfa6-4278-8748-473a2e7926dc · outbound

This paper cites In: CVPR (2024) 1, 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1, 3

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:25.052767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:18.725956Z digest=sha256:465fabc0db2c2a7dd180d17d506f454f14604532b4f353a40997b0da4d0aa2f5

Observation 81106527-a8cd-4c8f-a833-43e854334668 · outbound

This paper cites In: ICLR (2020) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2020) 6

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:24.698975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:18.776700Z digest=sha256:30b087eaa9533f75c3e19c67dffe762ac5396194f67f2af5c17b9e36cd1f3fa9

Observation 0beed7bf-d1b4-4ce8-8fe0-58119cbd0be3 · outbound

This paper cites In: ICLR (2024) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2024) 2

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:24.434302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:18.945497Z digest=sha256:fa3fbd581a6f13057eea38b72a64eb5531ba1e1b71b6d23986044889d05f1438

Observation 0231cead-6e59-41bb-b365-ccd9cc2b50bf · outbound

This paper cites In: ECCV (2016) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2016) 2

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:24.113166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:19.078605Z digest=sha256:10494d99fcd6f07f8f5a9797aef536a765f8cf65bc235eb0fb918e12ddbc8af2

Observation bc52b499-29e2-4fb4-a2ff-eff8f96974b8 · outbound

This paper cites In: CVPR (2023) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2023) 2

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:23.824932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:19.188805Z digest=sha256:5251f08f34999966d3146fe1a68adf9eb3fd365c1ca30f3c899d52acfc6fedfb

Observation 3726dca5-5270-440f-946b-40c29d849afe · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.313543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.313543Z digest=sha256:e9dbad5144cc8ed46d4b300083a34ad80eed63f800d77ff1d6315ced2abe0a6e

Observation 5c636545-9057-455b-86a3-c57df509d1d2 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.466169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.466169Z digest=sha256:05070fc32179e75991e8b7be3d4f1b120e9352dc16c87f3bc0313e9522340aac

Observation cf98a928-97fb-49f1-a8b2-0426a301e87f · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.548969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.548969Z digest=sha256:5f878fb73dfaf8e76293529b7fc26bac545859c95a1d1231af8787137b7a98ef

Observation bc289c9d-427d-4610-8baa-f295ed061a79 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.619551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.619551Z digest=sha256:d872abb2adff7c86e94d096f2d5a02ea551435789375c3722ae08c87a49e4a00

Observation e819ecc3-4df0-474b-b1df-8a031c12164b · outbound

This paper cites In: ICLR (2025) 5, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2025) 5, 6

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:23.450916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:19.736433Z digest=sha256:326505565879d84e83ed7aeca222a32c6378a965782d72bef8740740e018f327

Observation 4bc26048-147c-4687-835c-a95a3e75083a · outbound

This paper cites an unresolved cited work.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:38:23.105549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:19.814174Z digest=sha256:5058dd2f32fe643b3503b31cf5401d577afc5c34e190d21facf305070b0dcbbe

Observation 904869a4-664e-4648-b434-788a107d910f · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.891685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.891685Z digest=sha256:ab5d070d2a93473e28a03a2d04a6c6681da6701b50c74a9ea0aa7bdf48772f1e

Observation 17b80f1f-639d-4ed0-8004-9db7e39d7fca · outbound

This paper cites MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.968980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.968980Z digest=sha256:df3af28b597fc1ff59168d8b5bf42d482a0774d8dc272da04870ebf54c271542

Observation 1cc34e76-338a-432b-b6d5-791c97873d5a · outbound

This paper cites an unresolved cited work.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:38:22.784460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:20.040411Z digest=sha256:937664a4862dbbb72e4e9aacf6bbf43a0eb991ffafded9e9ea51cfb2a94a8c36

Observation 291812b2-9a35-4724-8333-e582821e6ec1 · outbound

This paper cites In: NeurIPS (2023) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2023) 2

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:22.525218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:20.130340Z digest=sha256:675d15e2bfac1a0a07e0aa6cf95a49211332179b775921993203d06ca741d093

Observation 3aa8ffb8-1d66-486b-9788-233ae4cc58c3 · outbound

This paper cites In: NeurIPS (2024) 1.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 1

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:22.186809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:20.238682Z digest=sha256:5ca4b4749f2f584f32333cb445cfbc83a5489f7713aec08b6ca22cba0155b5b7

Observation d69fd75a-9aa7-4bd6-928c-b428aaaf0e05 · outbound

This paper cites ViLLa: Video Reasoning Segmentation with Large Language Model.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.326601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.326601Z digest=sha256:08420de81ba7f8086bf1dde19df11e86c7e4b1d0ffee6389473dc825d5b21a2b

Observation e973c4df-534a-4661-8449-cf606e65c9f5 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MLVU: Benchmarking Multi-task Long Video Understanding

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.409072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.409072Z digest=sha256:7efa37a5609f333268e8c3177240484a62d9eef95454053ebbe73983b1b14bf4

Observation e8ea2821-4834-40c5-8833-fc92a11bd334 · outbound

This paper cites In: AAAI (2018) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: AAAI (2018) 2

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:21.736590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:20.527167Z digest=sha256:19c1af9828b081cb26b0404b770476dc06552d048b67d6db16433486ac921e9e

Observation 4144ffb2-608f-478b-b583-7ffa57eae19f · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.626312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.626312Z digest=sha256:15095407010016ce874a3cac9625a3c7500c93f0b4b2f4b422127d2779db91fa

Observation f7f5f831-62fd-475b-bc0d-3474d9e72514 · outbound

This paper cites VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.743420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.743420Z digest=sha256:4e65c2040f76ada277221af8790c0e6359d91d0698057fbea517cb0b0359ce4b

Pith citing papers

No inbound Pith citation observations are available.