Pith. sign in

Paper Citation Record · LEDGER

xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2410.16267.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.16267 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:32:53.658205Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T06:57:28.323422Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dae1cfa4-b820-47e8-a865-124eaad7325d · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:53.658205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:53.658205Z digest=sha256:20d03b8e6d8446c706c11acc756ab2782cf666fcd04a1bca980b3adc537d268d

Observation 53677ca5-1ac8-42b3-9335-e3510cf0ee6d · inbound

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding cites this paper.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.354802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.354802Z digest=sha256:a3c55c3be60b0fd5f29f2b965646d774b9797391ae122369d748d72bc92cd4ab

Observation 83a72fcd-4822-4217-91f4-a59a1769c279 · inbound

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs cites this paper.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.415433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.415433Z digest=sha256:9c4236b97b98ac52d410ebf3b1b0163482389bde63bea31897a5a4463ea8a41e

Observation 2f9abcff-b849-49cd-b4bc-18e75325fbd6 · inbound

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data cites this paper.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:17.487286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:17.487286Z digest=sha256:32a4a6a89e8194ec61336d364195c67ea4240fd6fe9853d612b3a2aaa45a3e97

Observation ac4e92f5-44c9-4494-bc87-5a20e2f7b837 · inbound

Tracing the Arrow of Time: Diagnosing Temporal Information Flow in Video-LLMs cites this paper.

Tracing the Arrow of Time: Diagnosing Temporal Information Flow in Video-LLMs xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:05:53.165996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T03:04:27.841522Z digest=sha256:0feb3e5c90a5c73f569703e08b83a2624245d241bce3e35af8508adf4f0478f2

Observation ed9d0d5d-9b5a-4980-a5e6-040ceb006084 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:26.761795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:13:21.487431Z digest=sha256:8853125cf49a68cf05f5e88a39cca407f49a3b7927232ea8f10030f7642ab4ff

Observation 57b75438-a683-4f00-af2d-872cc0916474 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:57:28.326463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T06:53:42.726350Z digest=sha256:99c445284352686e2bac44793ad8ecb1ec923db4b0135aedb2a16dfd873f3163

Observation b7a55875-f353-4b32-9ba8-3b1c1b284223 · inbound

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA cites this paper.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:44.272531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:44.272531Z digest=sha256:7a075789f8f3ccd48f0adf3b55469f85727e34b4172b1bbdb9cfa9a4732d3b13

Observation a8f4a119-5ef0-4545-87e6-4b6fc4e5ed8a · inbound

Efficient Tracking and Understanding Object Transformations cites this paper.

Efficient Tracking and Understanding Object Transformations xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T11:55:11.477753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:55:11.477753Z digest=sha256:84d9bdc64e1579392a81d0c862171418d95ecb396df5c9f0e48965d5dffe5d3b