Pith. sign in

Paper Citation Record · LEDGER

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding

As of 7 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2507.15569.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15569 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:34:44.512500Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f2aa759d-540a-4dd3-9b35-fcdbadc800c3 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.NeurIPS, 2022.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Flamingo: a visual language model for few-shot learning.NeurIPS, 2022

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.354775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.354775Z digest=sha256:9fd9ece257f562cbdaae7b71fc75871a3ff13a31ca018a1d10b040142e05164a

Observation f12f1706-f49e-452d-bde1-3937039bfe13 · outbound

This paper cites Vivit: A video vision transformer.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Vivit: A video vision transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.358353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.358353Z digest=sha256:04eb98c1fb799e08acc43446150cb3dbf555b75fc9410d1edb6443f628a745c9

Observation a810cc3f-b520-4e8b-837c-051629199850 · outbound

This paper cites Exploring Visual Prompts for Adapting Large-Scale Models.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Exploring Visual Prompts for Adapting Large-Scale Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.361412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.361412Z digest=sha256:2402cf9ede7b4c01e901e031a49fe6828158eda22a5f9b9ef90404486cfa4119

Observation 9204a1de-db76-4350-988f-6541a262e604 · outbound

This paper cites Relevant intrinsic feature enhancement network for few-shot semantic segmentation.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Relevant intrinsic feature enhancement network for few-shot semantic segmentation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.186177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.364929Z digest=sha256:c3f88ad2fc18d06cb2b09adf981e418908a6a01e3d09f8ce820e3c0442f8534f

Observation d902335b-31c9-4ce7-b95e-cd24b9314925 · outbound

This paper cites Cores: Orchestrating the dance of reasoning and seg- mentation.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Cores: Orchestrating the dance of reasoning and seg- mentation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.177037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.368268Z digest=sha256:ab79058d27f8bfa29ae54570f6040189dc769ae095d581afe4e263532dc1ec13

Observation 97a9d4ab-5612-42dc-9568-a85eafd9d064 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.371708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.371708Z digest=sha256:c119b8cbd2fcf0cd23d1577983e0cb67095b723c65de882ccde1ca40d4abeb29

Observation ed27e943-8da2-48a9-b119-bd9c8fb48b69 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Activitynet: A large-scale video benchmark for human activity understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.161693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.374741Z digest=sha256:074ec55d9c10511adaf492c8b2419c650b9fd5836ba17d6ff30d9efc0a7c9ce2

Observation 047a7701-6c38-4280-b268-a02799f653c6 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Quo vadis, action recognition? a new model and the kinetics dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.151912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.377884Z digest=sha256:f429561470cb501415b9bb5f0128fde89b843cfb97f5ca8e8f9d810e33691b7c

Observation 31550c44-a080-4788-aebf-c472fb184c73 · outbound

This paper cites Deep temporal linear encoding networks.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Deep temporal linear encoding networks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.142530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.380704Z digest=sha256:c964acbb24757b2b41a3b6a90f0ed4a41cd7102b5273d15cc7b860860328f85c

Observation 01a65432-7fe1-4d89-abb6-9290af5a52c7 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.383416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.383416Z digest=sha256:6b24d021903e49e219a14955b33fcc1e96ae9298525378648f28e9cc7bc1f5fe

Observation efca8d38-954e-475a-8066-dd34ff2a7bde · outbound

This paper cites Convolutional two-stream network fusion for video action recognition.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Convolutional two-stream network fusion for video action recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.131871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.386251Z digest=sha256:1b72654de682fb5cca6ae604cd6178b5497f6c8dd997a484cfa7c86f9441e657

Observation 06b216d4-bf76-415d-bfec-5b4821cd3b8e · outbound

This paper cites Slowfast networks for video recognition.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Slowfast networks for video recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.388931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.388931Z digest=sha256:1f2b6568c88b2319a172497f4d544879a624662df38966df36955459eb4ccfa5

Observation fe50a836-005b-4f99-9b63-5f8df3e74e80 · outbound

This paper cites A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.391857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.391857Z digest=sha256:7f8e0143c94bbc27f8becef7e1e051028010dc3181a6cbed06db3eb69a879dda

Observation 6c83cb25-e520-4df3-9bd8-c0d822c1f4a9 · outbound

This paper cites Lita: Language instructed temporal-localization assistant.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Lita: Language instructed temporal-localization assistant

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.115808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.394727Z digest=sha256:2093e4146ad1dbda3fc29848adade9e2672ac71bc781554ac19cd0d205755607

Observation 8758f006-c3e4-4eaf-bf76-4b97f6f07d39 · outbound

This paper cites Vi- sual prompt tuning.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Vi- sual prompt tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.397334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.397334Z digest=sha256:3d8ce498282304069c56683558bd2710256de40699aa9c8d82d89d40338fa579

Observation 76e16600-94a2-4da5-a431-b8651859cbbc · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.099303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.400058Z digest=sha256:3ca50c48d670634eb6fc7641b9d6e5f88c22b5003e6fdbe12bd58060b486cd2d

Observation 6b2d06ef-f871-479b-bf41-758ceadb3ae8 · outbound

This paper cites Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.402876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.402876Z digest=sha256:60062acbcc366f5243b485e76f2bc4bf1447cad5dfb15aadd89275b1e7b88778

Observation 6efde5e5-e056-4173-8d90-cf61eefb4bd5 · outbound

This paper cites Large-scale video classification with convolutional neural networks.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Large-scale video classification with convolutional neural networks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.089195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.405446Z digest=sha256:c3fb8eba8a5de4beff87d1ced69a3e72d20f5193450f5019136d65e4642b2d0c

Observation 886d5255-bf2c-432e-9ee2-bafee1ce1371 · outbound

This paper cites Maple: Multi-modal prompt learning.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Maple: Multi-modal prompt learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.077119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.407980Z digest=sha256:12bf3e8bf2c190670d6865c69f34433d186f8d9f5b060bc8b044904385a70147

Observation 86d87ae0-451a-468d-ba72-f19c6f13db1c · outbound

This paper cites An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.410539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.410539Z digest=sha256:675a7de512d0599e538d197e5c60120e4404c9285463bf3f0a428238b34081da

Observation 78686686-890a-4340-87c9-c72b035b5a40 · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding LISA: Reasoning Segmentation via Large Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.413798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.413798Z digest=sha256:045ad3e61281b4d5f8b8e38ebff7475516d0c658648962fb73e3735a0d05e01a

Observation e67cff02-1234-4df9-85db-504bce9b20d6 · outbound

This paper cites Deep local video feature for action recog- nition.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Deep local video feature for action recog- nition

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.067593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.416913Z digest=sha256:7677e3708b55afbd66c2fce0b5523567c6f0cba32d7f9c5b54602ceebdcf24bc

Observation cbe3b580-50c7-4eb6-9bc5-41608f5d5610 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.419656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.419656Z digest=sha256:696f31fd5f3f0a3d2abbb96d3c8cbcc34f8e69bc262901bb16caa3c40e3b300d

Observation 1fe9950b-56f2-4ab2-b4a8-f37eaedb4cb7 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.422636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.422636Z digest=sha256:0656b4cbe8f8275cba7eeda4c465ed8c1ce6452843cb41a3dcc3b58e8c9023b1

Observation 3906ca21-74a7-436b-9fba-b566579c0520 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.057513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.425250Z digest=sha256:19fdedb2b8c1c1b6e924605005f3afc97ffc9febe73953433b09cd618bf6a564

Observation 4f78f31d-0889-4c27-ab81-189ccb23a05e · outbound

This paper cites Tgif: A new dataset and benchmark on animated gif description.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Tgif: A new dataset and benchmark on animated gif description

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.048001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.427983Z digest=sha256:26400fee454420f388fd4aeafdf2962872b4f29e5d78eecdcc87ca50ae3ad991

Observation 5d3997aa-1479-4679-a207-0baac91a52ba · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Llama-vid: An image is worth 2 tokens in large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.038073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.430570Z digest=sha256:d4f88c7d31b46bce48de7c85842591f0fffcb7ac634a7e9670a61994acedca84

Observation 927376e1-c046-4e7c-9352-a6564ff8a61d · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.433210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.433210Z digest=sha256:d91c913a0fbe5c9372dc0e72a13c6c7760a81f5869231dcc25d3ff40b0c91907

Observation 42037cba-cf27-4151-b814-ad705284f57f · outbound

This paper cites Visual instruction tuning.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Visual instruction tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.436170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.436170Z digest=sha256:cbcea87697534806b35cf858538bcf526736b53c5750acebb5d204af8f1e7baf

Observation f149dbf1-c940-4306-ba9b-dcb387911ee1 · outbound

This paper cites Pre-train, prompt, and predict: A systematic survey of prompting methods in nat- ural language processing.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Pre-train, prompt, and predict: A systematic survey of prompting methods in nat- ural language processing

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.022210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.438941Z digest=sha256:910c0db2a1336a22f4849fa792b9fee4e2bd49216b92b8336c9347dcc4a6d48f

Observation f949c9ea-8c37-4d96-aab7-71637291e73b · outbound

This paper cites St-llm: Large language models are effective tem- poral learners.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding St-llm: Large language models are effective tem- poral learners

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.011156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.441777Z digest=sha256:ae894c16f3aae2e5f5f821dbb0d3008b17ce97ab8a00a75b143b6332c266d558

Observation 0a00f510-85b7-4864-9be1-0d8183088207 · outbound

This paper cites PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.444553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.444553Z digest=sha256:6f73bf79fd1f1723c4cff209f72a71db43232f0865e85297c9f19d22877dcde2

Observation 25f1e70a-1b36-485c-9ad4-bcbe0a679412 · outbound

This paper cites Hybrid-level instruction injection for video token com- pression in multi-modal large language models.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Hybrid-level instruction injection for video token com- pression in multi-modal large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.001920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.447387Z digest=sha256:59b08fd59a6a91d922f1c46fd49dd3acc5b57278ee776dca4e9fe953a59b0491

Observation 22f7bdd3-7318-42a1-9806-3829e0bcfd82 · outbound

This paper cites Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.450026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.450026Z digest=sha256:cc99e6a91bb47b899d8e79b91dc4742b25f62f719c0a8c8a14535bd6b7044c1e

Observation b074dd2b-5d0c-4053-9b0b-bcacb9f59593 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.992531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.453154Z digest=sha256:d4e4ca9d6aa0da9574cfa1fa3f4fd3849ca73cc1e6ab3fde3af0a89ebc4a71f9

Observation a26b246b-3b7a-4459-9e97-76a92169cba1 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Learning transferable visual models from natural language supervi- sion

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.982375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.456204Z digest=sha256:f7ec3679564507081a9325f5e5167fbc049d46b3811c88948cd3ec311104e83b

Observation 0b4c4c9b-ef73-4f35-a271-363c8546f5fe · outbound

This paper cites Mul- titask vision-language prompt tuning.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Mul- titask vision-language prompt tuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.971774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.458932Z digest=sha256:0c964e7b229b091cbdaa85e24fe16aa36bfb53239092df7a4924cff9ba81d968

Observation a007cafb-bd39-468d-850d-151dc2858a56 · outbound

This paper cites What does clip know about a red circle? vi- sual prompt engineering for vlms.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding What does clip know about a red circle? vi- sual prompt engineering for vlms

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.962147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.461665Z digest=sha256:e879cb66e8e3d898b6548dbd6e4526c6363799523298c06a2fbdc504921b8f98

Observation 0b9390b3-047e-40e2-885b-53296c9df236 · outbound

This paper cites Two-stream con- volutional networks for action recognition in videos.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Two-stream con- volutional networks for action recognition in videos

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.952584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.464208Z digest=sha256:b96af5662e15a100e2dfa65baa9e5f95790d15161e729b55416e622b57c46c04

Observation 8392e049-9274-45a9-a852-e6fb520459f0 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Moviechat: From dense token to sparse memory for long video understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.943100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.466804Z digest=sha256:790ccb1848094ba83bc725f56b811d97604439a685b54c3ca25c2d560f01494e

Observation 5339d259-49d9-44ea-833b-97d2c51089c3 · outbound

This paper cites Ufo: A unified approach to fine-grained visual perception via open- ended language interface.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Ufo: A unified approach to fine-grained visual perception via open- ended language interface

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.469414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.469414Z digest=sha256:9ab2ae3a5ed9c051b604f72a2325906d87892540a0fb636fa408de9702035b10

Observation f116e752-67e6-49a2-b634-1073b17f5477 · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Qwen2.5: A party of foundation models, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.932971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.472057Z digest=sha256:310143d44881f62196f4730fec77e0c93c63bcc2cbea24c072f9b5410348c94d

Observation 90ac6e1f-5dd4-4f91-a9d4-2e4a9e11e204 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Learning spatiotemporal features with 3d convolutional networks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.474572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.474572Z digest=sha256:92e3e776cbc9336e6c689ac42e9669e8a332662a994f950d407c6b9124116f57

Observation bf7cea43-2d1b-484f-8642-ab3d228a0e81 · outbound

This paper cites A closer look at spatiotemporal convolutions for action recognition.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding A closer look at spatiotemporal convolutions for action recognition

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.917281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.477334Z digest=sha256:1ca35ef72d46a8dfba7b3c038bb22b2ec2e891cc7c626f0e7a07ac3ad82393b4

Observation 562445aa-e6ce-488a-ac70-6da7fe1054ed · outbound

This paper cites Action recogni- tion with trajectory-pooled deep-convolutional descriptors.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Action recogni- tion with trajectory-pooled deep-convolutional descriptors

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.907511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.480128Z digest=sha256:3495fd562a2e77ab09d5762343d5958a21e02bef900f5b6b7c56947e6f045c32

Observation 1ac8883f-6190-43e9-9dd7-56414c619435 · outbound

This paper cites Temporal segment net- works: Towards good practices for deep action recognition.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Temporal segment net- works: Towards good practices for deep action recognition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.898103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.482802Z digest=sha256:43113a9ae30b5c4e05b69bd7a38898e060dc9955ba9c68483247d8b0efac475f

Observation 2d08469b-7499-4da5-9a0c-291c211afaf6 · outbound

This paper cites Deep learning for video classification and captioning.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Deep learning for video classification and captioning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.887994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.485388Z digest=sha256:3ef91f62aec45998e195719cb0b1bdd5bfaffaa563b9aff8abd51b0450fec247

Observation 04648ed4-742e-4fc7-a6dc-0c4f81a1dd1e · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Msr-vtt: A large video description dataset for bridging video and language

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.878270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.488052Z digest=sha256:d7a9a53435935970578ce974c0310954c1cc79719df1d94cb8d1682315c78419

Observation 62968b3d-c61b-44b4-b4b3-72671e84ee31 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.490784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.490784Z digest=sha256:8e23512ea263a9b1cc5291ec84382a20cd830c8f4c688e6a5ea3bdfe55df5075

Observation 39de8a1a-26a0-4798-a23c-7dd48455c721 · outbound

This paper cites Beyond short snippets: Deep networks for video classification.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Beyond short snippets: Deep networks for video classification

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.868532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.493641Z digest=sha256:452901ec98b006f88dbb3da362641e20debf6e50190934a64fd0d0dac559d0f8

Observation 624eca69-ece9-48ad-a3fe-6ee505763a9a · outbound

This paper cites Unified Vision and Language Prompt Learning.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Unified Vision and Language Prompt Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.496621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.496621Z digest=sha256:dfe5e47eda6775e0f8f40463f7c0baafae7216f0d190ee1abfea7d1a4c8abfa1

Observation 4e413645-f57d-425e-b8af-00f7ad24a065 · outbound

This paper cites Sigmoid loss for language image pre-training.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Sigmoid loss for language image pre-training

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.858675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.499401Z digest=sha256:75dcba5764937db65aa9097a40febba4aad90b9f30650565a1a133369b280ed8

Observation cadc5b69-4b8b-4b22-94c9-89979ec767a7 · outbound

This paper cites Real-time action recognition with enhanced motion vector cnns.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Real-time action recognition with enhanced motion vector cnns

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.848526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:34:44.502074Z digest=sha256:e8b4a45a89519dc4098d276346cef1506283514b9eb20745fd805adf05e1dca2

Observation 5bb3d3fb-30bd-42fc-8a91-94d8c352daf4 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.504615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.504615Z digest=sha256:3fe7f188b4f21c6b94e9adc6d25191c427fb0f0bc60253db2fde2c576dbf0efb

Observation 2e549f92-9770-44e1-8e79-718f53c45890 · outbound

This paper cites Conditional prompt learning for vision-language mod- els.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Conditional prompt learning for vision-language mod- els

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.507428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.507428Z digest=sha256:a8816a758fc6c48c40371555ad764ed00d7beb0a47b776ed0d20082f83091da7

Observation 3e308a2b-6870-4be2-8fd7-30105dfceff0 · outbound

This paper cites Learning to prompt for vision-language models.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Learning to prompt for vision-language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.510044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.510044Z digest=sha256:e94a45adce9120a48f17d45c6f65f8dee431d38cb728644a653ebcb14d5338f1

Observation 35094d80-6c56-4e63-8554-ee5219e9d0da · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.512500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.512500Z digest=sha256:4caed25088b78c64ddbe00e99474c5e6614f5d591ca439bf1db02d68a619393b

Pith citing papers

No inbound Pith citation observations are available.