Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T17:50:50.164729Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2508.15641.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T17:50:50.164729Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-07T09:12:13.460169Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T09:46:27.899181Z
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 86d4c77b-4a91-45b4-be7f-d6c03b50dd93 · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Vtimellm: Empower llm to grasp video moments,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aefeeac5-90d6-4422-91d1-defef698cfa2 · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Videotree: Adaptive tree-based video representation for llm reasoning on long videos,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 364333ec-be1b-490d-a1e3-cdf078de5905 · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f557661a-375b-4218-bfff-7d2641477f6e · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Flamingo: a visual language model for few-shot learning,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0039bdd4-8112-4d21-9d69-1396c1188891 · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2769dd94-fbfa-4156-8c93-84d1125e4928 · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a68319b-ab34-403f-824d-f917f278a170 · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a9a7e9ff-10cd-4e27-b89f-d12e6896098a · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 094dc931-5348-485d-9c67-e8218fcb5610 · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be38c7f1-2134-4f61-a717-636175509c2f · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Longvlm: Efficient long video understanding via large language models,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 627e59f1-c3ca-40f8-aa2d-712cb0a35502 · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Video summarization using denoising diffusion probabilistic model,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9577ab86-3d90-4017-a53a-bb418e43127c · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9154289-0194-427f-acab-353318a02020 · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Videoglamm: A large multimodal model for pixel-level visual grounding in videos,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e4090a78-d4c2-491d-ae9d-d464582bf38d · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding VACE: All-in-One Video Creation and Editing
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ffbb9f7-fe06-45a1-8331-7f38b3832e13 · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Wan: Open and Advanced Large-Scale Video Generative Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffe4b62a-7b5d-47bb-9c87-db8bf28ba84e · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Can i trust your answer? visually grounded video question answering,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5f7169c-70dc-4323-abb9-7951d3639879 · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Phi-4 Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c83433-9103-4c3f-a97f-767d35310425 · outbound
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5238bb53-07cb-4d26-88e8-8d378f7169a6 · inbound
YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.