Pith. sign in

Paper Citation Record · LEDGER

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs

As of 7 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 0 inbound Pith citation observations for arXiv:2507.10302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10302 v1

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:38:20.743420Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

92 of 92 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c201e185-eee5-49ee-a9b0-c2ca80d845f4 · outbound

This paper cites GPT-4 Technical Report.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:13.744309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:13.744309Z digest=sha256:18b375b6d09371c9a6c49e4b01087bb9d7d9f447d357904eb71e3fee610536d3

Observation 43efa36a-bb39-4e18-8419-7712a080ed51 · outbound

This paper cites In: NeurIPS (2022) 1, 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2022) 1, 2

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:13.800454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:13.800454Z digest=sha256:215b10d5afba9da48dd98f124c35e99e28edd6e9287aa523919ef956ca273e88

Observation 724bd4d6-c78a-45d0-a194-cfe9edf3479b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:13.887420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:13.887420Z digest=sha256:de79fcd775499fc4044b9b9ba64de948f456f4d628d91eaf3d787692bf039736

Observation 9f9a2c07-49b6-4862-a614-6beafbde6294 · outbound

This paper cites In: ICCV (2021) 2, 6, 9.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICCV (2021) 2, 6, 9

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:13.993701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:13.993701Z digest=sha256:e7e1a79c15e77264d588c231d5648c83446403d9d5504c93ac52257c7e1aee4f

Observation 607be67e-a3c6-44a2-8956-93aab90f7d4d · outbound

This paper cites In: NeurIPS (2020) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2020) 2

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.035353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.035353Z digest=sha256:0c29120a82b5f8797c8d6b37616b35703d9a72958401b398e92cb10de7a849b3

Observation 03173b9b-2712-4504-8da1-8ca05b4a22fe · outbound

This paper cites In: ECCV (2020) 4.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2020) 4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.105134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.105134Z digest=sha256:f1994572034b6c6ede0b3e948eae0fd49f95e7cda64be1a4ba5205d0d589809a

Observation 7836d8d4-84b7-47c8-a079-162823de13a0 · outbound

This paper cites In: CVPR (2024) 1, 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1, 2

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.175728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.175728Z digest=sha256:49c137972b16448188342d2a2e6575889c396cd002a408e5947227b534d0cbac

Observation 4f3e1cfc-98a5-41bd-9d90-15aea3596228 · outbound

This paper cites Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.260667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.260667Z digest=sha256:94e6dd5156d5a4e17d6d5a0eedeeeb521e405cba34b1a664f53d5cecb7bc9bf2

Observation 345d4ff9-0653-4faf-8052-39fb18424663 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.332299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.332299Z digest=sha256:e909dc6f30b1bd58b3f773cd57854250d1d7be07008c72f21b80939b3e7ca1cb

Observation 66fe3c21-9130-4585-a951-609915ea4b3a · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.390971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.390971Z digest=sha256:483c207d34879de3898f094b6d926f4a640f8627e26f53afaa89f714f93517f7

Observation 2da34952-7e06-42c6-ad0f-7a60dfe52fc0 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.432529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.432529Z digest=sha256:f666bd6fd85ae4062b3b926401e8a4c4ac4208d723c5429fd2fe9c24fb7504bf

Observation 86b558ce-f96b-4733-9f88-c59990bc0dd7 · outbound

This paper cites I see you, Batman!.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs I see you, Batman!

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.513288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.513288Z digest=sha256:12e0c86ca89eef7b1fa4aac0c8a2cb87463fd64bc18ffec74df16dc5c1dbf75d

Observation 5c4a722a-da2d-42d3-bc83-d5940e9691a8 · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.586386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.586386Z digest=sha256:babd40e8fe229176f381bb4bcd095c6aee3e514796228f283241d079f8711ce4

Observation 5498de92-18fa-4680-97ac-b93c70780507 · outbound

This paper cites In: CVPR (2024) 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 3

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.671358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.671358Z digest=sha256:5b5a815c2d6c101d0102c339674f0c0eccc000dfe170b94e6f14eed951827b43

Observation b7ae5fcc-f264-4131-9fb5-268430b92ae1 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.763233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.763233Z digest=sha256:571f68bdf820e53e4d938ac0fc90e75d04990e8b9a1c7f4903abf93bab5c9a71

Observation f0194408-f5ec-4a50-9cf6-5439a4022bcf · outbound

This paper cites an unresolved cited work.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.830921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.830921Z digest=sha256:f8caca1e04f086ab815538a2d28d16349b5112b3eff3469c0a19b9cfbeba91bb

Observation 1eef3aac-44ed-4ba3-bf11-518e7d03ccfa · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:14.924271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:14.924271Z digest=sha256:1404e74b1db49026b93954631bd1bd4c7fcd0daf671092e42c2f6f53cce4d681

Observation eefbd9ad-4aaf-413d-9aaa-49f560c5227e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.022435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.022435Z digest=sha256:9f79ff3b21a61312aa28aec31bee12f6b9cd3b09dc0f4a8482fcf2b4bcc0659e

Observation cb31d280-a509-4892-8e52-0cb174b0c121 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.079259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.079259Z digest=sha256:57f3fec7f7b4436e1bf409d61f9bd632b111b7fd574594f2fdcfe92b2527c1da

Observation 41c84606-e81a-430b-9fc2-b1a174403ae8 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs PaLM-E: An Embodied Multimodal Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.184382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.184382Z digest=sha256:53d2dac6290f5a622ef2c07b9d3fd0d39a3d20733186bdf596859ced9837707e

Observation 8cd5c9eb-8562-4209-a371-08b49ddf1272 · outbound

This paper cites The Llama 3 Herd of Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.276356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.276356Z digest=sha256:cfa2161c91187ca42f3afd960e4344edea08a4979f7971eea0ba54f245ad2d11

Observation b2823f80-b137-44c5-b1ac-0516804f4955 · outbound

This paper cites an unresolved cited work.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.378346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.378346Z digest=sha256:d6ead16d2e9ca01e683424bf3ea1139f6b5b75ca4002bb7e6d41376c5ac6be2d

Observation e3776ed1-c023-4dcb-b04b-50b221e334e1 · outbound

This paper cites In: CVPR (2023) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2023) 6

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.426339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.426339Z digest=sha256:ebff918b0754a0e914819a76b57cdf1582afe6194e3f844353c7a618bf8f3268

Observation 0ce230bb-953d-4a1e-8900-a855f7c56296 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.473346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.473346Z digest=sha256:15a5d4ea00e15de0a2aeb6f2ceb5148733d68df81cce4c3bfdab46cd1e50656f

Observation ff4062fc-f51c-4b6a-8b42-aeb20bd31974 · outbound

This paper cites Planting a SEED of Vision in Large Language Model.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Planting a SEED of Vision in Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.540771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.540771Z digest=sha256:0f332c57d3f5f2abae16ae63f1399f3c5a352fba46b0339fd04751c891f399b1

Observation 02d3d1dd-4791-464b-b24d-bc252e7e3746 · outbound

This paper cites In: ICCV (2017) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICCV (2017) 6

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.614673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.614673Z digest=sha256:afc99f4abcabf368ca71b7dfd6b23effe330eb3d8b2a42569125b7851423019a

Observation 1a15b683-b156-4d4c-9959-6d4ed0728984 · outbound

This paper cites In: ICLR (2022) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2022) 6

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.687789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.687789Z digest=sha256:b2b012189a543834ce297bd6edd4a59ff42861b2f6f7a9299b10264ef36415e3

Observation 607dafda-84ce-4743-9e65-46991d82fbb3 · outbound

This paper cites In: CVPR (2019) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2019) 2

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:31.066212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:15.781525Z digest=sha256:19ed7066f00868761dd535877b1a3e2bd9a70aefc2909dcd718b7886a9d12747

Observation e98d4342-9857-4619-b6b2-377dac9c3b79 · outbound

This paper cites Mistral 7B.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Mistral 7B

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:15.849245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:15.849245Z digest=sha256:853d1d75ac49d64c7f5323450b45e0278ec1351d1825ea8d77d71b037ab54f6d

Observation 2c7f589a-3863-4465-a417-2f691152cc5e · outbound

This paper cites In: CVPR (2024) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 6

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:30.848354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:15.963244Z digest=sha256:4eebc3f657a8a66653bbaab2f8379824953d407022186303bc69be2694706c9d

Observation 6cdc7be4-f0c6-42f2-8e3d-3274e708f59b · outbound

This paper cites The Kinetics Human Action Video Dataset.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs The Kinetics Human Action Video Dataset

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.022354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.022354Z digest=sha256:13d770262b0d06a82d1256c802a80b6d4dc53e33c2c76c2248de83450f2ff876

Observation 0959ef81-d36f-470e-82ac-a4902e94d8aa · outbound

This paper cites In: ECCV (2024) 1, 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 1, 2

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:30.531168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.087756Z digest=sha256:1fb76a9609e8d9b7712b9f06a829712eabeb7bc6ce1ff4dd0698515926130d1e

Observation 538bee68-24e6-4c16-ae51-3ecc023ed6df · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.116489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.116489Z digest=sha256:21227e32de6b768779cd5a87b0d215c151dd11661adb0456aeec42534189274f

Observation b348445d-e214-4a77-a551-d348e98fc3e9 · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.192030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.192030Z digest=sha256:a8e6072b6dc6bbeb5803f1d2edc1723cd9cde2275393e928c73b12efb9fef75b

Observation d59a25e4-0fa2-4725-9687-d6d4acfc89ff · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.241699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.241699Z digest=sha256:e2bc6fb47964f53593999efcfbc1011816d7db402113dfef69f6ad2349b9e535

Observation 4520737c-1b66-4709-bf54-2658a49e8343 · outbound

This paper cites In: ICML (2023) 1, 2, 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2023) 1, 2, 3

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:30.215814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.297011Z digest=sha256:0806aeb6d5231066577ed1b66a3c13ad3a884269e2b9e75a1b492458053ad02f

Observation 7542d54d-f7fc-4283-b69d-2ada832e5bca · outbound

This paper cites In: ICML (2022) 4, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2022) 4, 6

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.902598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.371511Z digest=sha256:82ba5135a2e8bca94699967f71517855aa424f862722dd7dc6514b7ad8c67561

Observation 7c86479a-1c60-4aa1-b402-216da35395b0 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VideoChat: Chat-Centric Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.430652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.430652Z digest=sha256:9b5b8fbe629c647112eb26ec47d367797f276b43af2c3016e445702cda58933f

Observation d8289b59-9e87-4593-b8c3-81e9643c4161 · outbound

This paper cites In: CVPR (2024) 1, 2, 5, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1, 2, 5, 6

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.657002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.489156Z digest=sha256:5560f6349f6029a1629c3e5c07c5fc6aba9387067a93c2a6d72ab7697843732b

Observation 442b2771-6305-4adf-85c8-671c77800522 · outbound

This paper cites TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T17:38:21.336547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.555603Z digest=sha256:84f170d2d837c7bdb824c87a2356d577343bdccd5326e5189ffdfa45575ad6d2

Observation 3243f840-cfe4-4a19-8f2d-3f79e8bfccb5 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.589632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.589632Z digest=sha256:e44e8dd8d054f9dde98ec4e110a26dcfb5a6e2f5884d6270d302ddeb18649ee9

Observation 321d4b0d-5c05-495f-80a2-096a7362fa8d · outbound

This paper cites In: ECCV (2024) 5, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 5, 6

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.498383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.624560Z digest=sha256:9ea2777d0f42cbdc459b3c9f9bc1683cc9a73fb962637a816f69b82171649297

Observation 32e494bc-9636-4bc8-9cc8-1b200e7e6b36 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.675424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.675424Z digest=sha256:91e3ab0c8f40791dad0b5b85143d22e578e9bba4fb7477b955aabec76f1d5262

Observation 883d0dd4-0754-4599-8909-105cbb6e0d87 · outbound

This paper cites In: ECCV (2014) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2014) 2

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.304790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.723028Z digest=sha256:177b08caeb2c96076c0887918ce8ca0fa6b5399fe4690fad288a15c2dad07bf1

Observation e629ad1a-c779-4479-8d24-7deba78bd179 · outbound

This paper cites In: CVPR (2024) 1.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:29.062686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.765941Z digest=sha256:f5e9227e71ada797437e63fb4465a5d59f1b38865ef8407ca02839c0e701e539

Observation 4a44d768-9562-4cc9-afbf-b67f35e7967f · outbound

This paper cites In: NeurIPS (2023) 1, 2, 3, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2023) 1, 2, 3, 6

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.846803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.846803Z digest=sha256:41e93be3ccad91fcce35c003a00b2a7dbe401c0198d246f002fb14e6bed6d224

Observation 81f964ad-84e6-4297-9683-bfdb99e9498c · outbound

This paper cites In: ECCV (2024) 5, 6, 9.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 5, 6, 9

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:28.764386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.889577Z digest=sha256:b0c4aa1e72af5640c998bd7a4bd06d8b3b5e1ce387216c10e94fc290e882afb2

Observation 8461f797-6b66-4aca-a0ad-f46468ac7f1f · outbound

This paper cites In: NeurIPS (2020) 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2020) 3

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:28.539521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.916656Z digest=sha256:a75be2e7cb4497be6bc71afb15f6820ee25764b1f83f334368cc436a754e30e1

Observation 7d6b7801-4119-411d-8ec9-75f68c7a6da9 · outbound

This paper cites In: ICML (2023) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2023) 2

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:28.318729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:16.966244Z digest=sha256:50664a9851979a8fd5bcf57a7945f229c08ee20e3e45422187aa4ec8976306b5

Observation 5bdbdf56-eabd-4041-96e7-58489e03b796 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.002664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.002664Z digest=sha256:bdad29317d14b30ee0cbe30311d52511865773ff4992020f8aa881c7126d1381

Observation 8e14157d-7c11-4d81-a775-cbeed9b6125d · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.072193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.072193Z digest=sha256:7bb94e011e6c464dd4cdc7fb25ad2a0be1b53927fd37dfc7f820c3b6293c9f7e

Observation 67b073d9-1d6a-4165-9ec4-b8d6bcc8e4ac · outbound

This paper cites In: NeurIPS (2023) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2023) 6

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:28.175772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:17.130992Z digest=sha256:7ac557bc44a95a72b4f8630360c5d95d1b3174eaf5160927f5529404c41625fd

Observation 386b79ff-461c-439c-8740-a77dfb464328 · outbound

This paper cites In: NeurIPS (2024) 1.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 1

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.903975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:17.161209Z digest=sha256:7bbe622e2197251be508ac79a92769c6b99d5973ca0284b1ab8c20123c518277

Observation 49a829dd-f3dd-4985-bef6-b4571903db11 · outbound

This paper cites In: NeurIPS (2024) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 6

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.677008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:17.208737Z digest=sha256:3810065d8739dc276d096c436dd285f44a6fb8fd851eb3caada70d4f172c72d4

Observation ee7759d0-0519-47ec-96a9-a2904642e3e4 · outbound

This paper cites In: ICCV (2015) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICCV (2015) 2

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.487305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:17.278269Z digest=sha256:7b2273cdffd66c8855a354ea531be1e42fc9d0e9ace00afb8b1b85d185b955a8

Observation 43c0b533-4ac5-4a04-9d05-714a5d021326 · outbound

This paper cites Streaming Long Video Understanding with Large Language Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Streaming Long Video Understanding with Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.320965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.320965Z digest=sha256:257140ec0f4de39aba5db0fc70afd046cccadbd83c6651a82b756cb2c77358dd

Observation 7233cdfa-b97e-478c-92fb-9754b71359eb · outbound

This paper cites In: ICML (2021) 4.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICML (2021) 4

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.263986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:17.439568Z digest=sha256:7720afe25432d16348176c5015ea259f47dc9360a8592ea46af5a6a9e3cd26ed

Observation b7ffccc3-a817-4768-aebb-b72c3e2920ec · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.535347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.535347Z digest=sha256:39df77061828224bb6938b5f5a917bf6f7567a633df06ab056a1aa8458e0af95

Observation 6a05af73-cc51-4fa8-a9fe-06b5697fa399 · outbound

This paper cites In: NeurIPS (2017) 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2017) 3

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:27.011795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:17.677509Z digest=sha256:e00d0ec21eb2e2933983f363ece96281606578bba378ebd759618db3e971a687

Observation 8d2417f6-8f7b-4124-ad1b-50160aa41901 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.726972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.726972Z digest=sha256:29a86e681efa1425523383e0aeccce716b143b5c0db8b0c28c8fd0c44097a3f0

Observation fbe2b09d-1871-496c-8ba4-195f40160540 · outbound

This paper cites In: NeurIPS (2024) 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 3

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:26.719355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:17.788180Z digest=sha256:2c17cb8216051bea05ca68bd5cd42c5732776b1dee2ce55912656eb8f2899c2d

Observation d0e2a374-33f4-4108-a157-64655cfd577d · outbound

This paper cites arXiv:2409.02889 (2024) 5.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs arXiv:2409.02889 (2024) 5

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:17.876527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:17.876527Z digest=sha256:cf99927ada6bfb4e507871f0a9beff8f87342823a88869378c15e52418d776bf

Observation a9e902cd-cb82-47cf-82d4-603c9f294408 · outbound

This paper cites In: ECCV (2024) 1, 2, 5, 6, 9.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2024) 1, 2, 5, 6, 9

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:26.539128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:17.935120Z digest=sha256:470290ae8fee0b4fe7c119e51c5a9bf6705757f6551380c8a9a476c98aea92f2

Observation 511b4426-8679-44f0-a05d-b86c191e34b7 · outbound

This paper cites In: NeurIPS (2021) 2, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2021) 2, 6

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:26.310005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:18.007029Z digest=sha256:22b29c77d6a45cca5d84af5978ba48b5189166f233ae70956aa66fedc047882c

Observation 0de2c905-712d-4d08-b347-fd91b045fc66 · outbound

This paper cites In: CVPR (2021) 6 12.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2021) 6 12

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:25.983174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:18.031278Z digest=sha256:6a690fb9a33b8c7f2ecee2796cbf6b4b74d26981e61d088b802e9a7f44e0ee1f

Observation f7d431bf-cb92-4640-a746-f6f2056a90d8 · outbound

This paper cites Slot-VLM: SlowFast Slots for Video-Language Modeling.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Slot-VLM: SlowFast Slots for Video-Language Modeling

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.120829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.120829Z digest=sha256:92515cbadb457b52820bf710aeedde3f32e5862fee06ac381ca07dc72212299b

Observation ebaffa45-cddf-4b24-9b6d-ca4799c54f16 · outbound

This paper cites In: CVPR (2016) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2016) 2

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:25.628447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:18.244765Z digest=sha256:447ff8750f02cbf2e074fd2e57839c65ed146d6397e25345819f83d361a16485

Observation 8da4eebb-4466-491b-b306-00581c9b7091 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.355450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.355450Z digest=sha256:4a5b255c2219a7d5f245112da1f6947954839a6b92f91104ef61db53ecf87092

Observation 34460155-d1c6-48db-85f6-23598d400956 · outbound

This paper cites RAL (2024) 1.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs RAL (2024) 1

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:25.341643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:18.449200Z digest=sha256:a319ea02e697cac74189793c5e8960046e36d10c19599092a57ae0051c946e95

Observation 7bf3ec7f-6b4a-4085-a110-40881a5299fd · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.564455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.564455Z digest=sha256:1993ac508c11332857a1af98f1289db64665b6704abc06f47d93ef95961192da

Observation e5b3b342-ae0a-44a5-b0b4-153c0da017fd · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.659223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.659223Z digest=sha256:aac845923b28c12d8b0c321813fc6a0c4b2bf9d2d552f5e7ae33c679c542ef58

Observation 8e0faf98-bfa6-4278-8748-473a2e7926dc · outbound

This paper cites In: CVPR (2024) 1, 3.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2024) 1, 3

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:25.052767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:18.725956Z digest=sha256:5bbf424a8ea0e79a4a65270570af6f71e7e50250426971dca91e53100e7c4b36

Observation 81106527-a8cd-4c8f-a833-43e854334668 · outbound

This paper cites In: ICLR (2020) 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2020) 6

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:24.698975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:18.776700Z digest=sha256:c674e08b25f87ae6c4bf629da424c10794327110c942d5d4f38a84330469c9ac

Observation 0beed7bf-d1b4-4ce8-8fe0-58119cbd0be3 · outbound

This paper cites In: ICLR (2024) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2024) 2

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:24.434302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:18.945497Z digest=sha256:bd22ccd1047792036e8a61a659361027f06f8215156a3681af825ac3aa5520a1

Observation 0231cead-6e59-41bb-b365-ccd9cc2b50bf · outbound

This paper cites In: ECCV (2016) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ECCV (2016) 2

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:24.113166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:19.078605Z digest=sha256:a7176ac3fb2044a6616c676c279e67ee861363b8a98cb77d535c7c62be7f511e

Observation bc52b499-29e2-4fb4-a2ff-eff8f96974b8 · outbound

This paper cites In: CVPR (2023) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: CVPR (2023) 2

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:23.824932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:19.188805Z digest=sha256:fea6d49781225bc3315c48a1c89fb71cfc8a01fdcf454f9317a7fe37d14c1bd3

Observation 3726dca5-5270-440f-946b-40c29d849afe · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.313543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.313543Z digest=sha256:7b440c3518c6b5451a461d0eb2ffd85b114346527f2801dabc7cd21cfa6b93ff

Observation 5c636545-9057-455b-86a3-c57df509d1d2 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.466169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.466169Z digest=sha256:9ab6d9c0dced5b52672002f3f17e1f7e82dfc7b4f94312ae959d5120685d6440

Observation cf98a928-97fb-49f1-a8b2-0426a301e87f · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.548969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.548969Z digest=sha256:95b972eb7ecc283e4c04414c032db7e2b49a954826f273de610a8815d402c388

Observation bc289c9d-427d-4610-8baa-f295ed061a79 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.619551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.619551Z digest=sha256:3a5107036224e4ae785da73fc8cd8cab0373f06e9c2a098d97bb3009fea155de

Observation e819ecc3-4df0-474b-b1df-8a031c12164b · outbound

This paper cites In: ICLR (2025) 5, 6.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: ICLR (2025) 5, 6

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:23.450916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:19.736433Z digest=sha256:ed81a209308d6c35892aa8bfc43d62de3419bae79a8753b14e9bae7900b0dc05

Observation 4bc26048-147c-4687-835c-a95a3e75083a · outbound

This paper cites an unresolved cited work.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:38:23.105549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:19.814174Z digest=sha256:4950d65d89b9f84395f6940443172eb7e52067b857241fda5a3260534c92d6e1

Observation 904869a4-664e-4648-b434-788a107d910f · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.891685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.891685Z digest=sha256:5c2322f4d43228a8bed86797643249c48003f8e6435d3c44cce6ca04e51bf67d

Observation 17b80f1f-639d-4ed0-8004-9db7e39d7fca · outbound

This paper cites MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.968980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.968980Z digest=sha256:f4d2ec93a7d9ad532da87324c7ff5cb49aff4ebb44aab156d8f1863a0fe3fb67

Observation 1cc34e76-338a-432b-b6d5-791c97873d5a · outbound

This paper cites an unresolved cited work.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:38:22.784460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:20.040411Z digest=sha256:5b1433db06f5a9375e10f760f4f02647ea84eff10f38904974ea5240c8779297

Observation 291812b2-9a35-4724-8333-e582821e6ec1 · outbound

This paper cites In: NeurIPS (2023) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2023) 2

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:22.525218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:20.130340Z digest=sha256:f9c32eec99788d5d751339dcbf518c6fafd7f3a5be0a7c91bffb43212d0b7948

Observation 3aa8ffb8-1d66-486b-9788-233ae4cc58c3 · outbound

This paper cites In: NeurIPS (2024) 1.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: NeurIPS (2024) 1

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:22.186809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:20.238682Z digest=sha256:9a4af9190475f2162d9a254adcb33ff99caa08d2f3169dd6c30fe46b1de500ca

Observation d69fd75a-9aa7-4bd6-928c-b428aaaf0e05 · outbound

This paper cites ViLLa: Video Reasoning Segmentation with Large Language Model.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.326601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.326601Z digest=sha256:32ac0d8c5eca6adb4a0e441572257197ad7718b536aa5e6741678417d528850d

Observation e973c4df-534a-4661-8449-cf606e65c9f5 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MLVU: Benchmarking Multi-task Long Video Understanding

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.409072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.409072Z digest=sha256:7f19ec6faed7c7c50e52d894da2702f2c8bdf45754a6eefda3f41e3ba0fcf2c5

Observation e8ea2821-4834-40c5-8833-fc92a11bd334 · outbound

This paper cites In: AAAI (2018) 2.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs In: AAAI (2018) 2

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:38:21.736590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:38:20.527167Z digest=sha256:f697622c55aeccfd76dd0a3a0086104787485c7b9a42e2fcf16adc74196e69a6

Observation 4144ffb2-608f-478b-b583-7ffa57eae19f · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.626312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.626312Z digest=sha256:a458437a9c06220fa529cbab192aaee93d5b621033d2b63d8ec4c5306ad4fe51

Observation f7f5f831-62fd-475b-bc0d-3474d9e72514 · outbound

This paper cites VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.743420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.743420Z digest=sha256:33c1e2badc08f473b0027da5f4060a55e054e076ac1718295d385e7de44fa1e4

Pith citing papers

No inbound Pith citation observations are available.