Pith. sign in

Paper Citation Record · LEDGER

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding

As of 5 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2508.15641.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.15641 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:50:50.164729Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-07T09:12:13.460169Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:46:27.899181Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 86d4c77b-4a91-45b4-be7f-d6c03b50dd93 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Vtimellm: Empower llm to grasp video moments,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:52.029281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T17:50:48.479947Z digest=sha256:34a3a773ae31213a4d5b6e393ce2392793690ae6f9c1fca3c4e9b56d86501cf3

Observation aefeeac5-90d6-4422-91d1-defef698cfa2 · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Videotree: Adaptive tree-based video representation for llm reasoning on long videos,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.781116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T17:50:48.511981Z digest=sha256:cd1e4a0c5b45c83fdeb1b0d5c369716d0ab44fa75d9017be70ac7bcf14db3af6

Observation 364333ec-be1b-490d-a1e3-cdf078de5905 · outbound

This paper cites Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.583649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T17:50:48.594059Z digest=sha256:ecaa413fb0f7510c1a9fc4c8ea2408ba053a144943ced451d09b0552b2d11179

Observation f557661a-375b-4218-bfff-7d2641477f6e · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Flamingo: a visual language model for few-shot learning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.377592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T17:50:48.678859Z digest=sha256:b94fbfb69acc701b3ce9fc7ae418994f001a04e0f2cceda94f878de9cc4938fe

Observation 0039bdd4-8112-4d21-9d69-1396c1188891 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:48.760200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:48.760200Z digest=sha256:50f3f11952ce6569e4fa1966d9543f5a0167d1bedc6e8b8243bb697143b5b9ad

Observation 2769dd94-fbfa-4156-8c93-84d1125e4928 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:48.860892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:48.860892Z digest=sha256:cd992bf9728d9d9a6e11b7273e6e35ea9d936eb5e14ec9dfb9589a3f4ef1e52f

Observation 4a68319b-ab34-403f-824d-f917f278a170 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.190776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T17:50:48.962891Z digest=sha256:6f3899bb13a5587b5040bc565f5d6c4030966fddf426d63eddf80bba5dd629da

Observation a9a7e9ff-10cd-4e27-b89f-d12e6896098a · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.044330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.044330Z digest=sha256:e1fba107232c05ea5afa222e3f66e7b8d9a7ab3c58232f026ec70a4e4abf76c2

Observation 094dc931-5348-485d-9c67-e8218fcb5610 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.173346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.173346Z digest=sha256:1f3865c0cab88a506b3ec10c665a2528feb167bc6f2c0bf668b1bcf7ab5c70c4

Observation be38c7f1-2134-4f61-a717-636175509c2f · outbound

This paper cites Longvlm: Efficient long video understanding via large language models,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Longvlm: Efficient long video understanding via large language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.011850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T17:50:49.247131Z digest=sha256:3cfd33a5a2852e829c2f4b86a5a9e47f70f07e374c21c4c1022df6b9d2db38d8

Observation 627e59f1-c3ca-40f8-aa2d-712cb0a35502 · outbound

This paper cites Video summarization using denoising diffusion probabilistic model,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Video summarization using denoising diffusion probabilistic model,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:50.847614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T17:50:49.390336Z digest=sha256:f3852f4542a342fd31b5ebaaea29dfa73a39facc41a7d451f1e6324cbedef7a9

Observation 9577ab86-3d90-4017-a53a-bb418e43127c · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.465440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.465440Z digest=sha256:07b6eafa89df482dced24d9d73c4d09e32e7e951ea92d804bf95a1a929d16fe9

Observation f9154289-0194-427f-acab-353318a02020 · outbound

This paper cites Videoglamm: A large multimodal model for pixel-level visual grounding in videos,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Videoglamm: A large multimodal model for pixel-level visual grounding in videos,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:50.709653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T17:50:49.612907Z digest=sha256:f877108b8ac35154ac33c9605e0cf7db7469ed4dc341e5bb613d6500300a0db3

Observation e4090a78-d4c2-491d-ae9d-d464582bf38d · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding VACE: All-in-One Video Creation and Editing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.717507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.717507Z digest=sha256:fdee7ee7c6af79dc24d4a6af4f72cf8610bbf8ce4cad58209b7408bd5cf29a1c

Observation 8ffbb9f7-fe06-45a1-8331-7f38b3832e13 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Wan: Open and Advanced Large-Scale Video Generative Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.870293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.870293Z digest=sha256:4d44b0aeb68d468c94e8af9d52958853cd7e440220e831f330393b321cb550e8

Observation ffe4b62a-7b5d-47bb-9c87-db8bf28ba84e · outbound

This paper cites Can i trust your answer? visually grounded video question answering,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Can i trust your answer? visually grounded video question answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:50.460962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T17:50:50.011389Z digest=sha256:421510dc078c5303033d0d84ef822f45d7d5807f74a784726bfdaf625281880b

Observation c5f7169c-70dc-4323-abb9-7951d3639879 · outbound

This paper cites Phi-4 Technical Report.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Phi-4 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:50.085949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:50.085949Z digest=sha256:86447f620f9f4778ec679cdcfed022270c5497998c5f5c2570cef43cd50e9b57

Observation 26c83433-9103-4c3f-a97f-767d35310425 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:50.164729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:50.164729Z digest=sha256:34ea35e0e93bcd1d42acd8c6b10f5ec3d64b02843b878a78c202d973d9152b2e

Pith citing papers

Observation 5238bb53-07cb-4d26-88e8-8d378f7169a6 · inbound

YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal cites this paper.

YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:46:27.901967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T09:12:13.460169Z digest=sha256:a0a04c88968d981e7ff147290ec800534f97643b7e22c5ee9b4059f054460ab0