Pith. sign in

Paper Citation Record · LEDGER

Vidi: Large Multimodal Models for Video Understanding and Editing

As of 24 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 8 inbound Pith citation observations for arXiv:2504.15681.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15681 v3

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:25:33.395157Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:03:32.341621Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T20:46:13.651443Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 75545c5d-9008-4c3c-bb3d-62c7e814ba34 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Vidi: Large Multimodal Models for Video Understanding and Editing Gemini: A Family of Highly Capable Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.212021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.212021Z digest=sha256:50cd01c4eba4061a528147f883c3211bfe2ed528501a78bc7d2d6edb64ecd5a6

Observation c1b128ed-282b-4a18-be43-1223ecefd36d · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Vidi: Large Multimodal Models for Video Understanding and Editing Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.217824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.217824Z digest=sha256:20f9ea225bed8e427d9c5f84f9e1e619bab2c6bed99e641790e157f940c1ba29

Observation 1dcb4d13-fd59-4cc3-98b2-9741a3fb604d · outbound

This paper cites Qwen2.5-VL Technical Report.

Vidi: Large Multimodal Models for Video Understanding and Editing Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.224325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.224325Z digest=sha256:0ce2834273d72f6ea6a5de26163be515f792620c3fcf22e49eb82b4a9f44aaaa

Observation 908c11c5-134e-461f-9765-87df8cc8216b · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Vidi: Large Multimodal Models for Video Understanding and Editing InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.230301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.230301Z digest=sha256:d209502cc5830dceea4dab1e37eba0e889d5f6f0cfd75313ec217559684acf95

Observation a55b34b9-08aa-4fbb-b82a-9c9a86b8f33a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Vidi: Large Multimodal Models for Video Understanding and Editing Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.235593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.235593Z digest=sha256:26957df7e5e0cd3a9a28c488062ae1f901396a147a1a08030335a02a130042aa

Observation cfdf50ae-f463-4b04-ac97-3d186445b46a · outbound

This paper cites TALL: temporal activity localization via language query.

Vidi: Large Multimodal Models for Video Understanding and Editing TALL: temporal activity localization via language query

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:34.042280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:25:33.240648Z digest=sha256:723f8a6fd5c45cffc8ef2e80eacc517a52f9a0434448289e953b58bb515aec66

Observation 214ad4d7-1925-4596-82c5-de054e189fc1 · outbound

This paper cites LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos.

Vidi: Large Multimodal Models for Video Understanding and Editing LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.246113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.246113Z digest=sha256:9bf9e5d96680f3f91a3420eb9295dfab4e737ed41308ec9a1312ac385b028635

Observation 2f40f9da-8876-4edc-b47d-0853101dd3e1 · outbound

This paper cites Fullstop: Multi- lingual deep models for punctuation prediction.

Vidi: Large Multimodal Models for Video Understanding and Editing Fullstop: Multi- lingual deep models for punctuation prediction

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:34.026104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:25:33.251244Z digest=sha256:6d89f3b9be3353490bd73514edd017b56523710bff9e82b894b7c9804cc06318

Observation eca0cda8-904b-41e6-bd67-3ee360580d5f · outbound

This paper cites an unresolved cited work.

Vidi: Large Multimodal Models for Video Understanding and Editing Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:25:34.011321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:25:33.256100Z digest=sha256:2ec074b96f05c34aa3203b275d734cb1d5a4346e1b1d7036fc8c801735f34187

Observation 2f08ae5d-be3a-4036-8e2e-b5338d041836 · outbound

This paper cites GPT-4o System Card.

Vidi: Large Multimodal Models for Video Understanding and Editing GPT-4o System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.260694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.260694Z digest=sha256:f3f905d354208bf8cbbe6f3ffe99347b0dfba6e19e10be1b4feaeef7453e0bb2

Observation 993668c6-c7f5-494d-b87e-68b41cfece4f · outbound

This paper cites Mistral 7B.

Vidi: Large Multimodal Models for Video Understanding and Editing Mistral 7B

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.266005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.266005Z digest=sha256:50ce1273fc34dd2e24670f536cbb0231121bfd3ab203d9262120232f44548903

Observation 925c9294-5948-4462-937f-d5de7c897a1f · outbound

This paper cites Dense-captioning events in videos.

Vidi: Large Multimodal Models for Video Understanding and Editing Dense-captioning events in videos

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.996400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:25:33.270683Z digest=sha256:b4e977e1ef5a9b36b9986f758f4cd332c2be7af30e42034331fa3e3c59395aa6

Observation 52eb6fbe-0c18-4296-879c-85f36cf4b945 · outbound

This paper cites D-Attn: Decomposed Attention for Large Vision-and-Language Models.

Vidi: Large Multimodal Models for Video Understanding and Editing D-Attn: Decomposed Attention for Large Vision-and-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.275272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.275272Z digest=sha256:407b001c8ba5da07e6c1545c68681c81a0f564dcd5eb6fb4ab8c8f37a860dbf8

Observation c4de6e23-b1fc-47d4-86dc-8339a6c0d366 · outbound

This paper cites Berg, and Mohit Bansal.

Vidi: Large Multimodal Models for Video Understanding and Editing Berg, and Mohit Bansal

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.980908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:25:33.280295Z digest=sha256:2dd49d7a851716514691769618d83d411fc12df24bbbc93c77c58b5b7409e068

Observation 921dc973-69cd-4219-bec8-e91e21bac565 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Vidi: Large Multimodal Models for Video Understanding and Editing LLaVA-OneVision: Easy Visual Task Transfer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.284877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.284877Z digest=sha256:ffe159423bee0f33ef51a1bbdeec3d7f7892fe8ad2de32ee10c0f5d0bb353bb8

Observation 79db2b3c-7825-41f7-b54c-cabe13a8bc0e · outbound

This paper cites Structured Context Transformer for Generic Event Boundary Detection.

Vidi: Large Multimodal Models for Video Understanding and Editing Structured Context Transformer for Generic Event Boundary Detection

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.289574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.289574Z digest=sha256:872d4105cf1d0de9be12b502e3ac38c9831687614954a4a806b07c3a242a220a

Observation ef0b583b-9242-4a78-a9a7-563b4c57b1f7 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Vidi: Large Multimodal Models for Video Understanding and Editing LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.294219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.294219Z digest=sha256:c83779c424850001a563f2f394f708b610fadd3bd3fa76ebedb8eb8d9958d86b

Observation 5688aec5-1346-4d95-b5d5-46b36a739e0c · outbound

This paper cites an unresolved cited work.

Vidi: Large Multimodal Models for Video Understanding and Editing Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:25:33.965757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:25:33.299785Z digest=sha256:f824ef043f533759443607fbb49c867face525b60353e7a4aa9331d217fc7da8

Observation cc2472a3-33a7-4874-8443-8ee7660d6d0c · outbound

This paper cites Visual instruction tuning.

Vidi: Large Multimodal Models for Video Understanding and Editing Visual instruction tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.950194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:25:33.305127Z digest=sha256:31b6255fc674f9fd9fe9334cb927d3ca8932dca8ac35e91d1364448999ffeab6

Observation 4212afa9-c0d8-412e-92ff-3a793e0328f8 · outbound

This paper cites Decoupled weight decay regularization.

Vidi: Large Multimodal Models for Video Understanding and Editing Decoupled weight decay regularization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.933844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:25:33.309832Z digest=sha256:5db8ba1d0076536e6ba105da57ad72dc803c7780e8e740f635337c146908d43d

Observation 7e6fb3ee-1f7a-4dc7-9853-8e70650fb3ec · outbound

This paper cites ZoomV: Temporal Zoom-in for Efficient Long Video Understanding.

Vidi: Large Multimodal Models for Video Understanding and Editing ZoomV: Temporal Zoom-in for Efficient Long Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.315744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.315744Z digest=sha256:6c8544fdbf6cca6122f7511c0b5f4a2f99e86ef32256b0af99b57e335175a1a9

Observation aced7a25-51c5-4384-9cf3-48c9aa8dc898 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Vidi: Large Multimodal Models for Video Understanding and Editing Robust speech recognition via large-scale weak supervision

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.913817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:25:33.320936Z digest=sha256:581275406dee50cf8a831f05ae88167d030d7f9d1a14ea9af4e4aed6d64ad629

Observation 03452171-a7be-4592-a52e-f95ebc2b79f9 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

Vidi: Large Multimodal Models for Video Understanding and Editing CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.326194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.326194Z digest=sha256:29fb084df2ff60b17e91a38f73adccd14a25f46bfdd37d74c487944d577f9462

Observation 175a2ce4-8450-4f99-8052-409ba80863a3 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Vidi: Large Multimodal Models for Video Understanding and Editing Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.893528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:25:33.331326Z digest=sha256:4f6923822a10999832e66ab2a4991d7766dbb2bc520e197eedb40ad8ab0bbbf5

Observation 4feb0279-91c2-4237-ab43-02099d89dfe7 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Vidi: Large Multimodal Models for Video Understanding and Editing Gemma 2: Improving Open Language Models at a Practical Size

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.336090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.336090Z digest=sha256:3b27c0c4f1b689a1afb76184a9982fd443440a1af19ecfb39828035be209f43f

Observation 550b9d51-acc6-4c81-ab00-9b987a757ac1 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Vidi: Large Multimodal Models for Video Understanding and Editing Moviechat: From dense token to sparse memory for long video understanding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.875255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:25:33.340709Z digest=sha256:d7e7fcfa0b507d850792131376de9660e1b4c3be84c9195cb31c9afddf33f88a

Observation dbe16d8b-ee4a-4e05-ac19-985e8cd3a226 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Vidi: Large Multimodal Models for Video Understanding and Editing SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.345897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.345897Z digest=sha256:7dc78fab789ef34fa54b4ae5b5169e6a9ad914d895dda5f76bf368022d53e9be

Observation 3a0225ed-6fde-4a55-91fd-4881953bc472 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Vidi: Large Multimodal Models for Video Understanding and Editing Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.858125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:25:33.351356Z digest=sha256:6a55cf9c93f00328bbd14899a3e896f568c7927746fdd6d71af8a797a9e46845

Observation 5edaf4e0-9884-47d4-bbe6-5eda2d6cf1f4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Vidi: Large Multimodal Models for Video Understanding and Editing Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.356094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.356094Z digest=sha256:e756231dbe8670e1a2aa99a972b15ce8977e7fa95cf783d28da2ccaa2d5dc99a

Observation 654cb160-705f-458e-947d-6be61b1082fa · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

Vidi: Large Multimodal Models for Video Understanding and Editing Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.360996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.360996Z digest=sha256:0c19159f5b8b3a6b358b2db74a4f6c2acfedbd7bde810ca79b3bd87a4ddf8245

Observation a6c0f6e6-1ee6-4bd0-85f4-c9cdd88e94e5 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Vidi: Large Multimodal Models for Video Understanding and Editing LVBench: An Extreme Long Video Understanding Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.365977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.365977Z digest=sha256:1bd77c8ae04b0b44c50dae8fe82c5e7137a99140c9614d11ed3616e99b0415b5

Observation c5303013-c728-4eb4-9848-818e6ea3fa1e · outbound

This paper cites Chi, Quoc V.

Vidi: Large Multimodal Models for Video Understanding and Editing Chi, Quoc V

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.841343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:25:33.370899Z digest=sha256:d89e3ec2048d876bd13b3c8937f606dfd3f34b84dc91d58fc76def1dd6cf87d1

Observation e7b47912-3644-4c6b-96a0-0d3c51f5c5be · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Vidi: Large Multimodal Models for Video Understanding and Editing Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.825454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:25:33.375805Z digest=sha256:5baf02d64451b611b9ecbb2af1bbba77f40234155805459376fa7753dcd252be

Observation 3774260e-0ffa-4758-8876-f4d5e768af80 · outbound

This paper cites T*: Re-thinking Temporal Search for Long-Form Video Understanding.

Vidi: Large Multimodal Models for Video Understanding and Editing T*: Re-thinking Temporal Search for Long-Form Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.380253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.380253Z digest=sha256:952213758e3804d7fac96116daaba7d2d60ec4e8cf4f3bf7b7ab598ea7c7e1b9

Observation 21125145-8afe-4382-a0dd-731c852c6066 · outbound

This paper cites Sigmoid loss for language image pre-training.

Vidi: Large Multimodal Models for Video Understanding and Editing Sigmoid loss for language image pre-training

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:25:33.809062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:25:33.385653Z digest=sha256:b6123e49ef9af589829ed76c48078f3632b9f5845b3848d291bf361166a1bd8a

Observation cce8b1d4-b89a-4d8d-933d-e1a496a1b9a1 · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

Vidi: Large Multimodal Models for Video Understanding and Editing InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.390194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.390194Z digest=sha256:f05ece2aefb5b7fdd79cfaf020270527d793c23caf8324c8713912f37170b442

Observation a27b6c80-bc29-42e8-9a38-94194eee2fb4 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Vidi: Large Multimodal Models for Video Understanding and Editing InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:33.395157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:33.395157Z digest=sha256:026ba595582363868786c5ae7612053894b37997fcb1851096017105efae0285

Pith citing papers

Observation c784aa37-b1ec-47ec-b4c2-37215622e784 · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:32.341621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:32.341621Z digest=sha256:cb6334ca17e04ddd28037fc5043b96bed114098dda0aaf69134c68b9b7ffe461

Observation e1dcda03-a4ba-4173-8a0b-7819012ba321 · inbound

Timeripple: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space cites this paper.

Timeripple: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T22:12:45.376792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:12:45.376792Z digest=sha256:adfbdd7789013f7c035265b0f8f5054b90ea6e43d8aade8e8ccac07a474099fa

Observation 60d8b032-c162-4b77-b4cf-53ec4b9aae5d · inbound

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation cites this paper.

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T13:21:10.692684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:21:10.692684Z digest=sha256:9e09b88b161db8d4ef5362b9bb6caa9bef82909f041b683f0830a7b8ee2d9963

Observation 3895ec79-b6ea-41d3-bc29-d9331292d0d8 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.126091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:95746e52cb57ab22402d7292d80043f1aebc3a039585291fa7d5a662832b3e83

Observation 20247188-fc73-4b75-8c2c-6c629235a58c · inbound

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning cites this paper.

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:13.660350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T08:04:15.238840Z digest=sha256:9f6d19e77347ab243a719530434b58d6759bfd1d93e7a14a3bff72b61876cd54

Observation c6fcef02-2efa-4065-b867-a466e360b740 · inbound

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning cites this paper.

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T06:04:16.637378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:04:16.637378Z digest=sha256:4cae9e5be51913b90364f1dd833c7b9e84e26282019e50ed368a211e6db9fd05

Observation 8536bac6-8a77-49a4-be09-1afb22f38f08 · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:47.844379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:47.844379Z digest=sha256:b84ce21e1e1f5b683e8952c278c9774695564f6885030ec34ba31c16c23dece1

Observation 48bf1810-df25-4868-a715-14b523634f5d · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:15.986807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:15.986807Z digest=sha256:2b02483ec02d853666a1d6d0eeef525a1eddd6acc36a05d30562024bfb474740