Pith. sign in

Paper Citation Record · LEDGER

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding

As of 22 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2508.15641.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.15641 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:50:50.164729Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-07T09:12:13.460169Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:46:27.899181Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 86d4c77b-4a91-45b4-be7f-d6c03b50dd93 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Vtimellm: Empower llm to grasp video moments,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:52.029281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T17:50:48.479947Z digest=sha256:eca8ba479745889c1499dd5fc28f863ebaed2db38d725eacf689d15f57aa4528

Observation aefeeac5-90d6-4422-91d1-defef698cfa2 · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Videotree: Adaptive tree-based video representation for llm reasoning on long videos,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.781116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T17:50:48.511981Z digest=sha256:5971b28e0c920c827c02c2da5835bd7fcbb08a0a9f2d106440c3af6ffcc83443

Observation 364333ec-be1b-490d-a1e3-cdf078de5905 · outbound

This paper cites Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.583649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T17:50:48.594059Z digest=sha256:9472f4aea122b9508e9c04d3901dff52f27cee660e112cf2be104f5a4babbdf6

Observation f557661a-375b-4218-bfff-7d2641477f6e · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Flamingo: a visual language model for few-shot learning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.377592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T17:50:48.678859Z digest=sha256:09d35cfe5d6eeaafd5aac21dad19af7f5200aeec94d001f62b70f1b6eec83f19

Observation 0039bdd4-8112-4d21-9d69-1396c1188891 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:48.760200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:48.760200Z digest=sha256:a52fe16999b058963b12af8d6e7f1bf16bab09684753d945b36a84c86f99b8a1

Observation 2769dd94-fbfa-4156-8c93-84d1125e4928 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:48.860892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:48.860892Z digest=sha256:9d5eadfb125a10a4da77550289567bb072591ea587194b0add38efa30b964768

Observation 4a68319b-ab34-403f-824d-f917f278a170 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.190776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T17:50:48.962891Z digest=sha256:f827160b291d3413794109d05f057ccfd2ac2510b56a623b8ba3c931d2fe9dbc

Observation a9a7e9ff-10cd-4e27-b89f-d12e6896098a · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.044330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.044330Z digest=sha256:2d4b68b47cf748fff38004763c491f70a5b920fb7f16bf7f3b870996ad473223

Observation 094dc931-5348-485d-9c67-e8218fcb5610 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.173346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.173346Z digest=sha256:f0be5c92d8a29e7dfe1e44eff4439f2586cf4a073ba8f0015c3697c64667c23c

Observation be38c7f1-2134-4f61-a717-636175509c2f · outbound

This paper cites Longvlm: Efficient long video understanding via large language models,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Longvlm: Efficient long video understanding via large language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.011850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T17:50:49.247131Z digest=sha256:e6ae3b03d2d9520a4dea47b81bc2a03421bf2d0652d24c706c419cdcd8f97894

Observation 627e59f1-c3ca-40f8-aa2d-712cb0a35502 · outbound

This paper cites Video summarization using denoising diffusion probabilistic model,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Video summarization using denoising diffusion probabilistic model,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:50.847614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T17:50:49.390336Z digest=sha256:9344de2914b821624315d99d58b700e77058f3166555f79f367bbb29cfc3e6f4

Observation 9577ab86-3d90-4017-a53a-bb418e43127c · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.465440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.465440Z digest=sha256:8cf39bda50b7b1750fbd731da60695799a08e7637159655709d993e70c471a89

Observation f9154289-0194-427f-acab-353318a02020 · outbound

This paper cites Videoglamm: A large multimodal model for pixel-level visual grounding in videos,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Videoglamm: A large multimodal model for pixel-level visual grounding in videos,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:50.709653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T17:50:49.612907Z digest=sha256:1fd92fb63beb078955ebf639e58e5b799d97e73f0ce6dbec6d82bbc3ea511703

Observation e4090a78-d4c2-491d-ae9d-d464582bf38d · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding VACE: All-in-One Video Creation and Editing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.717507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.717507Z digest=sha256:6f0371ec811e601177c0b664d0eede25579e3e5dbecd62f55f6ba4cc57491111

Observation 8ffbb9f7-fe06-45a1-8331-7f38b3832e13 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Wan: Open and Advanced Large-Scale Video Generative Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.870293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.870293Z digest=sha256:ebf0d875c134c8f0ef06f18d07597f51dd9881c260dc04813c995a367f0935e7

Observation ffe4b62a-7b5d-47bb-9c87-db8bf28ba84e · outbound

This paper cites Can i trust your answer? visually grounded video question answering,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Can i trust your answer? visually grounded video question answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:50.460962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T17:50:50.011389Z digest=sha256:a35caaf09422504108eafb27eec65e3d3a3d154240492679a70dceac8f8f9261

Observation c5f7169c-70dc-4323-abb9-7951d3639879 · outbound

This paper cites Phi-4 Technical Report.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Phi-4 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:50.085949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:50.085949Z digest=sha256:2a5457fb4de6c32aab019acc5faec04ff7a7c48313a5677bf7faf90948aac94a

Observation 26c83433-9103-4c3f-a97f-767d35310425 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:50.164729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:50.164729Z digest=sha256:bb20c5ecf60181dddc454279dfcb8c2a29f31b05c77cacf2c6b3512bc9fac6d2

Pith citing papers

Observation 5238bb53-07cb-4d26-88e8-8d378f7169a6 · inbound

YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal cites this paper.

YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:46:27.901967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T09:12:13.460169Z digest=sha256:8487abfa2d8ca3a24aa532386532fce25a164435cf913bc2e08096094077c3e7