Pith. sign in

Paper Citation Record · LEDGER

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding

As of 17 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 3 inbound Pith citation observations for arXiv:2411.14401.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14401 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:17:19.535023Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T00:19:26.153682Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4775f871-409b-495c-ad49-c996aa385278 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.482163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.311077Z digest=sha256:c051ab33c25e344df6e76f880b1ee2b5b6f8e2b6ec62c2f63528794040046d1e

Observation 9894c35c-047e-4413-9367-df150fa01062 · outbound

This paper cites Token merging: Your ViT but faster.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Token merging: Your ViT but faster

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.316997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.316997Z digest=sha256:21607d47889585593910c13ee5cddae42cdf69999b7690294416e3fe8980e3c6

Observation 65cb87c9-ba34-4c34-ad75-5da0bc51ba73 · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Collecting highly parallel data for paraphrase evaluation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.454974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.321812Z digest=sha256:e20649c2073f6dcad1a0a52750d8c173fa1b9c887b82abd8235aa3f89880188c

Observation df3283eb-4378-402f-ac2e-3d494be2ce4c · outbound

This paper cites Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.326998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.326998Z digest=sha256:70ad57d497f820b38776107af68c75859850291671f31022caf42d36bd0efbaa

Observation 0df4e6c6-1ce9-4d70-8453-9f03fbf54617 · outbound

This paper cites HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.332297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.332297Z digest=sha256:08049c837458ca91bf19af71336a4b40be7a3a2f2874972b0006dd8be6d41786

Observation 522e2b5f-218d-4e71-a728-a82a41ccd501 · outbound

This paper cites VideoLLaMA 2: Advanc- ing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding VideoLLaMA 2: Advanc- ing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.439033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.337077Z digest=sha256:ebb7ea44c245989510e646c918ad3993f44e8a01476d08569574d20d93a79a61

Observation 6584ecec-8fcd-411b-a0b5-d29c20b22968 · outbound

This paper cites Video-MME: The First- Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-MME: The First- Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.422508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.341685Z digest=sha256:f00afb395eb4b50705092712647f573b8e9d2ff78ff905d9f1d65cb4fb1dbd65

Observation e97becb7-a974-43ac-8878-56c3eda69ca5 · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.346184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.346184Z digest=sha256:8a04f1c858db47e7e9fec635ccbf6a54bf705dbb5e2bc5b7bdc603fdb76153cc

Observation d13b6b4d-4bfe-4311-b0fd-4d9bbd12c0f1 · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.404545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.351596Z digest=sha256:caba20cd8ba96aa5706fe85fbce2b879ca4136e76cb2482291d47527348f5836

Observation 34f17add-0dea-426e-8c3b-acf3a602a184 · outbound

This paper cites Evaluating Open-Domain Question Answering in the Era of Large Language Models.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Evaluating Open-Domain Question Answering in the Era of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.356837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.356837Z digest=sha256:c65ad4bcffb42bd0f8e48872ee7252d37ece2a0038816da89616559d59ceb5de

Observation 94a946b4-2d98-48ea-8c81-090b8683b43c · outbound

This paper cites An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.382311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.362862Z digest=sha256:7fca379e5b23f964a425dea0bb5eeccca7df15a5c10e01bd76fb24764cfce4ac

Observation b368254a-fb71-477d-99c1-1bb6d90eb9d5 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.367588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.367588Z digest=sha256:b05f046e42ae29c933c72eaddaadfcb6d6deb90e05ce8a01602c98fc14002c97

Observation 126393c9-28ef-420b-9a0b-77f28d11f8cc · outbound

This paper cites Inten- tqa: Context-aware video intent reasoning.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Inten- tqa: Context-aware video intent reasoning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.360658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.373377Z digest=sha256:82989c1944e954d58ee36e5071b4d5a732fad9bf58284a95230d292e8a0dd8ec

Observation 1fb44f72-8ff2-45f3-89fb-3b0efa5067e9 · outbound

This paper cites Uniformer: Unified transformer for efficient spatiotemporal representation learning, 2022.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Uniformer: Unified transformer for efficient spatiotemporal representation learning, 2022

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.335870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.379498Z digest=sha256:6d8e1e76508123600d0b1cc59edc33f09608194d3ab00fa0e6e35e37dd7e288a

Observation 45b39613-a40d-4e1c-becf-74ee4e5336e2 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding VideoChat: Chat-Centric Video Understanding, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.321606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.384344Z digest=sha256:34f44592061e05eab17ed46f8e557373df2e43f04fd9397c7295d554696e370c

Observation c968a893-9c6c-45af-9fa5-73e2e717b92b · outbound

This paper cites MVBench: A Comprehensive Multi- modal Video Understanding Benchmark, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding MVBench: A Comprehensive Multi- modal Video Understanding Benchmark, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.302761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.388360Z digest=sha256:20f052aabd68dcf9ad4bd1e221ef59c842b09ff99cef1f72264075a1ac344ba0

Observation f95d33bb-a7a4-421d-9b07-971045cc1870 · outbound

This paper cites Tgif: A new dataset and benchmark on animated gif description.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Tgif: A new dataset and benchmark on animated gif description

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.285900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.392754Z digest=sha256:5dbf17a55ba243b23d2a5b9d49e017c3e5795c01aa074a501f19d31fa245acdd

Observation 40103667-a67b-4cca-955c-0a043a37b599 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models, 2023.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.269655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.396866Z digest=sha256:ea4f0b22ae9abe12401eee3c38f7fed92478c6eb578fd84a33a87ec14e9965d5

Observation 2f2d01cb-2286-4b52-9849-91ba195fdb1c · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Llama-vid: An image is worth 2 tokens in large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.254266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.400928Z digest=sha256:224b92e69ac4c3372c7e8da44187188f5d441a488a3b8bd978bf8cb87af8bc25

Observation 2af3598e-f90c-4b78-9ef6-6f98a5f579b8 · outbound

This paper cites Video-LLaV A: Learning United Visual Representation by Alignment Before Projection, 2023.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-LLaV A: Learning United Visual Representation by Alignment Before Projection, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.237582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.405590Z digest=sha256:cb44b8ceb393bde834bc344e1f3281cb342416c8b910acde616417460c786f70

Observation 27d28afd-2203-4e27-afe6-7f9115ebeb05 · outbound

This paper cites Tsm: Temporal shift module for efficient video understanding.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Tsm: Temporal shift module for efficient video understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.217520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.409359Z digest=sha256:6e1ae9fefed70f8c2b4e76e65745aba8fe8a40503ec5eabd4d390fc3ca15e1a5

Observation be1c8c6a-6027-42fc-b188-e631b843d082 · outbound

This paper cites Video swin transformer.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video swin transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.413389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.413389Z digest=sha256:e5e354c06ef0b96c8afa5eaabcc4f53c956dfa44fe394675ae23b3adb926e1f1

Observation c5bb445c-ae61-4b73-bc79-99f772d81188 · outbound

This paper cites Vista-LLaMA: Reliable Video Narrator via Equal Distance to Visual Tokens, 2023.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Vista-LLaMA: Reliable Video Narrator via Equal Distance to Visual Tokens, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.188660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.417067Z digest=sha256:8040c5714d63940c061ba4de58ae5006079f4b4768c2084ff2c709c3dd71af90

Observation 292b679e-7e22-440e-9a36-3bc0a2e54d21 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.420733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.420733Z digest=sha256:bf57ff18fc036a1a5f0d3a21c41c3bf95e971cf1f3cfc970cd584d44e46f52bf

Observation 9c2212ef-61a1-4c2c-9eeb-a7056a73a074 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Mod- els, 2023.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Mod- els, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.170409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.424893Z digest=sha256:40937123e57ac07c822ceedd80ef4154e71a0b40e23ab6164d87e2fc843a415a

Observation a76213b4-79b5-476e-805c-d43f79e92590 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.156272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.428685Z digest=sha256:9262dd6ce0f88cbe7dc643f1d3c5b5d164fe47d62c913961efac090ec6c1e49e

Observation 09290a55-b8f4-42ff-82c4-3002a6a7c6e5 · outbound

This paper cites Foundation Models for Video Understanding: A Survey.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Foundation Models for Video Understanding: A Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.433014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.433014Z digest=sha256:f55d3a04bc17255437d28c65d34e33cda65f69f331e7bf1d12f0192c24913996

Observation 00d517ab-25a8-4355-8797-38d79fdd9fc8 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- 9 form video language understanding.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Egoschema: A diagnostic benchmark for very long- 9 form video language understanding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.139469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.437718Z digest=sha256:431fd378aaf5bcd7d7ed69ff5ffab3cf1adc5056feee2d397b8a1dc6f862f753

Observation 3ec775d1-81f7-4664-bf67-6c7f25c5f6df · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.441536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.441536Z digest=sha256:ccd71928579dec048eef8eec232d0ebba8b3653ff29659471e1229b3c9751e85

Observation 698362b6-d6f1-4625-83ea-a10dffb579ac · outbound

This paper cites Less is more: Pay less attention in vision transform- ers.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Less is more: Pay less attention in vision transform- ers

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.123968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.445865Z digest=sha256:dcecc0c19f1b4debe806fa734c6eaf0ded82a60169c48e079f793abdd7a85cbd

Observation 8ba4bb6d-d006-4869-a91d-94c5e1a9bccb · outbound

This paper cites Effi- cient parameter-free clustering using first neighbor relations.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Effi- cient parameter-free clustering using first neighbor relations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.107949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.449954Z digest=sha256:de2326ff52c76600da3ef0eb296e8411b4d2f05eaf562bbf51c36d09d398efc7

Observation bd5e2a6f-3fa4-4272-8a7c-5b678a64b268 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.094030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.454068Z digest=sha256:5d406eaa3a870cf0d8418e2094ed22d7c1b04e7574104ae79b420399977de56d

Observation dd73bd8a-9856-4504-a44a-673f2a94970c · outbound

This paper cites Video understanding with large language models: A survey.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video understanding with large language models: A survey

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.457831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.457831Z digest=sha256:3827145eb72b81cd2269ba3ed3438d2aa3561d624b925caa2753664ef60eba77

Observation d753d32a-3dbd-45bb-9e13-9d78255acdce · outbound

This paper cites Internvideo2: Scaling foundation models for mul- timodal video understanding.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Internvideo2: Scaling foundation models for mul- timodal video understanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.079942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.462316Z digest=sha256:c8a6b6d70ba218c0a3720dce4fa03e0a71f715a299dcd58f3aaa6313c062e144

Observation db43c629-b1c1-4ccf-a52b-a567266de72e · outbound

This paper cites Freeva: Offline mllm as training-free video assistant.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Freeva: Offline mllm as training-free video assistant

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.064393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.466380Z digest=sha256:b63b5655fe6774c416efb5717f7e11809e798bdfa83ac69a211112083a5e6021

Observation 2bed5b56-0624-48f6-aa5b-886f14588489 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Next-qa: Next phase of question-answering to explaining temporal actions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.047090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.470673Z digest=sha256:2c34c7327d2d2d64ba2d3b325aa4032f6340d208318805592e72dbe6cbb4fa66

Observation 0d929b18-6a74-46fd-a4ee-455e44dc4ca0 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Msr-vtt: A large video description dataset for bridging video and language

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.027795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.476488Z digest=sha256:dab58101f4006ca680e6e3c0121a36f5201bd9f6fdaf8e52bd8766c379c455ad

Observation 14923024-3d8d-418f-8728-03e3bee1f7c7 · outbound

This paper cites PLLaV A : Parameter-free LLaV A Extension from Images to Videos for Video Dense Captioning, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding PLLaV A : Parameter-free LLaV A Extension from Images to Videos for Video Dense Captioning, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:20.010444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.481713Z digest=sha256:de88b8602476a4fa7ee3bc2f227dab5e9fef53695c2d6d95d96be026a30630d6

Observation 7399f2ba-5ace-47e2-b8b8-63840eade38a · outbound

This paper cites SlowFast-LLaV A: A Strong Training-Free Baseline for Video Large Language Models, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding SlowFast-LLaV A: A Strong Training-Free Baseline for Video Large Language Models, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:19.982895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.485762Z digest=sha256:e03fbb821274041ceeb50b038cf74cfc730ab47edc10093778533c86d837705d

Observation 69bd3635-6fdf-4f89-b31f-2dbe3a87845e · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Self-chained image-language model for video localization and question answering

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:19.964391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.493075Z digest=sha256:f68d05913e97351ae90f6559e567272f2712d90ed331656fd5e2d6fc6e4cb71f

Observation 76c52acc-6394-407b-9ea5-ec27461b406d · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Self-chained image-language model for video localization and question answering

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:19.946256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.498409Z digest=sha256:fea1727a262793ef910a68741ea6efa9dc32bead96f922387aef7a8c8b133fa1

Observation ed617feb-2b64-4461-a4a3-8a9f2adf14dd · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:19.930947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.502684Z digest=sha256:88f65c5365fc65028570a5dafb205796b9b091b47f3ed03bf4452c83a5a1de64

Observation e2cf9dd5-bb5c-4c4a-8943-43c299e5ecbe · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, 2023.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:19.916189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.507722Z digest=sha256:ce99dff74efa4a327d09e8c9d1c5c4553ed972ef79501e7da2a01cbb24418626

Observation 84dfa728-6de5-460e-b61f-942252e29a55 · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.513134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.513134Z digest=sha256:2c5593cbb08a8a1f6e3af7402b5cbda4004e89a91946889393fa0fe874b4111b

Observation 4ade763f-3bac-4b4c-987a-fe95caf799b4 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Llava- next: A strong zero-shot video understanding model, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.517778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.517778Z digest=sha256:93984dfe748b97d223471dfac2d85dc5fa5e030286e7cb28e87e5136c0fd29cf

Observation caba4785-ef58-4e5d-91a3-f8d139b4d05f · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Llava- next: A strong zero-shot video understanding model, 2024

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.522013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.522013Z digest=sha256:803061f46dc39af6d9cad71eee46c0f74528d95c554520974a338478a551e41a

Observation c64262c1-0a17-4e9b-be81-f3fbb8d16758 · outbound

This paper cites RankCLIP: Ranking-Consistent Language-Image Pretraining.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding RankCLIP: Ranking-Consistent Language-Image Pretraining

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.526139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.526139Z digest=sha256:76621d8146d36126bab715c29dad4d38a36772ce8f1f57cd71d64e6cea9bb70d

Observation 46dc1631-eb8a-425b-998d-de6f1d6bcb56 · outbound

This paper cites Multimodal guidance network for missing- modality inference in content moderation.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding Multimodal guidance network for missing- modality inference in content moderation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:17:19.884148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:17:19.530734Z digest=sha256:980cbc480fc5e3362807495654248fe1b7c476025801f2057ab1272815cc667f

Observation 23d25d77-fef8-4ec5-8ece-71083439c966 · outbound

This paper cites A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming.

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:19.535023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:19.535023Z digest=sha256:225e43ab4f59d548adcc0695375cf26a6125761e7b4c27fa9baac0119d20e002

Pith citing papers

Observation 3eea6c55-aa84-4112-8e27-05f5de688bc8 · inbound

Mosaic: Cross-Modal Clustering for Efficient Video Understanding cites this paper.

Mosaic: Cross-Modal Clustering for Efficient Video Understanding Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:10:34.350836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T16:07:36.133404Z digest=sha256:55a94186edccd3cb22622be96b627fbdaec77284fb92616c9fe59e6c3102ac96

Observation 01acc121-5f1e-419a-bd97-6bb4d4ed9dca · inbound

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models cites this paper.

GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:17:50.839649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T19:15:13.205594Z digest=sha256:d3b32d525b57019e4cacfe03410620ba483aed271b8df926082451222d25e796

Observation adcf274d-f2a2-4a79-bf6c-9297ffb9eab4 · inbound

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding cites this paper.

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:58.352126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T00:19:26.153682Z digest=sha256:b489fb215b3cccff8dc0109e57711d4097d9e430c5274cce05185a8257615c17