Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:34:44.512500Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2507.15569.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:34:44.512500Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f2aa759d-540a-4dd3-9b35-fcdbadc800c3 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Flamingo: a visual language model for few-shot learning.NeurIPS, 2022
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f12f1706-f49e-452d-bde1-3937039bfe13 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Vivit: A video vision transformer
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a810cc3f-b520-4e8b-837c-051629199850 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Exploring Visual Prompts for Adapting Large-Scale Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9204a1de-db76-4350-988f-6541a262e604 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Relevant intrinsic feature enhancement network for few-shot semantic segmentation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d902335b-31c9-4ce7-b95e-cd24b9314925 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Cores: Orchestrating the dance of reasoning and seg- mentation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 97a9d4ab-5612-42dc-9568-a85eafd9d064 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Is space-time attention all you need for video understanding? In ICML, page 4, 2021
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed27e943-8da2-48a9-b119-bd9c8fb48b69 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Activitynet: A large-scale video benchmark for human activity understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 047a7701-6c38-4280-b268-a02799f653c6 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Quo vadis, action recognition? a new model and the kinetics dataset
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 31550c44-a080-4788-aebf-c472fb184c73 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Deep temporal linear encoding networks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 01a65432-7fe1-4d89-abb6-9290af5a52c7 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efca8d38-954e-475a-8066-dd34ff2a7bde · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Convolutional two-stream network fusion for video action recognition
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 06b216d4-bf76-415d-bfec-5b4821cd3b8e · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Slowfast networks for video recognition
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe50a836-005b-4f99-9b63-5f8df3e74e80 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c83cb25-e520-4df3-9bd8-c0d822c1f4a9 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Lita: Language instructed temporal-localization assistant
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8758f006-c3e4-4eaf-bf76-4b97f6f07d39 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Vi- sual prompt tuning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76e16600-94a2-4da5-a431-b8651859cbbc · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6b2d06ef-f871-479b-bf41-758ceadb3ae8 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6efde5e5-e056-4173-8d90-cf61eefb4bd5 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Large-scale video classification with convolutional neural networks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 886d5255-bf2c-432e-9ee2-bafee1ce1371 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Maple: Multi-modal prompt learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 86d87ae0-451a-468d-ba72-f19c6f13db1c · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78686686-890a-4340-87c9-c72b035b5a40 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding LISA: Reasoning Segmentation via Large Language Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e67cff02-1234-4df9-85db-504bce9b20d6 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Deep local video feature for action recog- nition
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cbe3b580-50c7-4eb6-9bc5-41608f5d5610 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fe9950b-56f2-4ab2-b4a8-f37eaedb4cb7 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3906ca21-74a7-436b-9fba-b566579c0520 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4f78f31d-0889-4c27-ab81-189ccb23a05e · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Tgif: A new dataset and benchmark on animated gif description
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5d3997aa-1479-4679-a207-0baac91a52ba · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Llama-vid: An image is worth 2 tokens in large language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 927376e1-c046-4e7c-9352-a6564ff8a61d · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42037cba-cf27-4151-b814-ad705284f57f · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Visual instruction tuning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f149dbf1-c940-4306-ba9b-dcb387911ee1 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Pre-train, prompt, and predict: A systematic survey of prompting methods in nat- ural language processing
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f949c9ea-8c37-4d96-aab7-71637291e73b · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding St-llm: Large language models are effective tem- poral learners
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0a00f510-85b7-4864-9be1-0d8183088207 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25f1e70a-1b36-485c-9ad4-bcbe0a679412 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Hybrid-level instruction injection for video token com- pression in multi-modal large language models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 22f7bdd3-7318-42a1-9806-3829e0bcfd82 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b074dd2b-5d0c-4053-9b0b-bcacb9f59593 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a26b246b-3b7a-4459-9e97-76a92169cba1 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Learning transferable visual models from natural language supervi- sion
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0b4c4c9b-ef73-4f35-a271-363c8546f5fe · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Mul- titask vision-language prompt tuning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a007cafb-bd39-468d-850d-151dc2858a56 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding What does clip know about a red circle? vi- sual prompt engineering for vlms
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0b9390b3-047e-40e2-885b-53296c9df236 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Two-stream con- volutional networks for action recognition in videos
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8392e049-9274-45a9-a852-e6fb520459f0 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Moviechat: From dense token to sparse memory for long video understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5339d259-49d9-44ea-833b-97d2c51089c3 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Ufo: A unified approach to fine-grained visual perception via open- ended language interface
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f116e752-67e6-49a2-b634-1073b17f5477 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Qwen2.5: A party of foundation models, 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 90ac6e1f-5dd4-4f91-a9d4-2e4a9e11e204 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Learning spatiotemporal features with 3d convolutional networks
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf7cea43-2d1b-484f-8642-ab3d228a0e81 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding A closer look at spatiotemporal convolutions for action recognition
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 562445aa-e6ce-488a-ac70-6da7fe1054ed · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Action recogni- tion with trajectory-pooled deep-convolutional descriptors
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1ac8883f-6190-43e9-9dd7-56414c619435 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Temporal segment net- works: Towards good practices for deep action recognition
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2d08469b-7499-4da5-9a0c-291c211afaf6 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Deep learning for video classification and captioning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 04648ed4-742e-4fc7-a6dc-0c4f81a1dd1e · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Msr-vtt: A large video description dataset for bridging video and language
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 62968b3d-c61b-44b4-b4b3-72671e84ee31 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39de8a1a-26a0-4798-a23c-7dd48455c721 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Beyond short snippets: Deep networks for video classification
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 624eca69-ece9-48ad-a3fe-6ee505763a9a · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Unified Vision and Language Prompt Learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e413645-f57d-425e-b8af-00f7ad24a065 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Sigmoid loss for language image pre-training
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cadc5b69-4b8b-4b22-94c9-89979ec767a7 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Real-time action recognition with enhanced motion vector cnns
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5bb3d3fb-30bd-42fc-8a91-94d8c352daf4 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e549f92-9770-44e1-8e79-718f53c45890 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Conditional prompt learning for vision-language mod- els
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e308a2b-6870-4be2-8fd7-30105dfceff0 · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Learning to prompt for vision-language models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35094d80-6c56-4e63-8554-ee5219e9d0da · outbound
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.