Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:01:21.032754Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2507.17050.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:01:21.032754Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T12:13:38.580739Z
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation eba5d90d-1e6b-4e03-803b-235051e41324 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Optimizing marketing strategy: a video analysis approach
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3086b80a-7c54-4933-b054-130272eabcec · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f63121dd-c975-455f-a460-2afcb70cbe12 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Mm- au:towards multimodal understanding of advertisement videos
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 117d45b5-b834-4480-99bb-a65dc24eff5f · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Blog: E-commerce product videos with examples, 2025
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 727b0d93-9bae-42ff-bee5-0021c6fa1fa8 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f87acd8-6b4d-4b05-8db7-23d0659e1b2e · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5ad6c06-6171-41f0-a505-2ca5f57f794a · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Yolo-world: Real-time open- vocabulary object detection
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b71f51b6-6e6b-4d98-ab16-c62e5e6ae583 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 958d8f7b-fda2-4189-b8d7-be08f4085284 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Deepseek-r1: Incentivizing reasoning capa- bility in llms via reinforcement learning, 2025
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 612a4483-0884-4de4-ba48-09f93e430382 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60b03cf5-3ed3-41be-a131-eb254f80051d · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Sketch, ground, and refine: Top-down dense video caption- ing
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b451ef04-ecc6-4ebf-9417-15dd9c8e62e9 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Video-mme: The first-ever compre- hensive evaluation benchmark of multi-modal llms in video analysis
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d38fc442-b07b-49da-ab2f-5ce2082f2202 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abe00bd2-dd7f-454b-aa97-67645fbf4b99 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models MiniCPM: Unveiling the potential of small language models with scalable training strategies
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5adda0b6-68c2-4ad5-a57d-064b04934878 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Lita: Language instructed temporal-localization assistant
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7265ce93-0b36-4462-9633-cd40092fb2a0 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Automatic understanding of image and video advertisements
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2f3b36da-67d6-4e87-aaa5-961b35cebf3f · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models A better use of audio-visual cues: Dense video captioning with bi-modal transformer
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f95af62b-679a-4149-a605-ae30d0115040 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Multi-modal dense video captioning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d56f9c4f-cbfe-41c8-90e9-57f9c5b9c2de · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Video re- cap: Recursive captioning of hour-long videos
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 053435b6-fb70-4682-bfb0-3a5909386f8b · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Drive targeted marketing in retail with video ana- lytics insights, 2022
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 16221969-ee3e-4131-ad09-1d1227d2a7ed · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Dense-captioning events in videos
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2c9170b1-69de-47f5-8039-3b8022545839 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Llama-vid: An image is worth 2 tokens in large language models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4d10611-336c-473e-b65e-c957ca7e16e1 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Development and challenges of object detection: A survey
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4a99b57e-5102-42ef-87f4-15f547d55e79 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2b000a7-c7a4-4477-9912-9fb4940a4473 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efacaea5-5a1c-49bb-a7b8-7814f71bca7a · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ca176c6-65d6-473b-9254-22bbb358aa98 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Dense video captioning: A survey of techniques, datasets and eval- uation protocols
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 01df04a6-cf88-41fd-94cc-01ab46d00a6f · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Timechat: A time-sensitive multimodal large lan- guage model for long video understanding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6e6dc0b6-d181-47ce-b5d2-eec2163aa5e3 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Moviechat: From dense token to sparse memory for long video understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66e6dad2-3916-4c63-b8ca-151b68520dbd · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Video understand- ing with large language models: A survey, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 09ebf718-9af9-4545-a0fb-85318b5503e2 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Llama: Open and efficient foundation lan- guage models, 2023
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 430ded0d-8fda-4b06-a5c0-369f4fd5360e · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Cider: Consensus-based image description evalua- tion
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c38a615f-9637-465c-956a-0d0a52776827 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4476f321-8a75-4258-8440-923db242c431 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 46b28dad-b922-49d2-972b-75cbb8de2139 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models End-to-end dense video captioning with parallel decoding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0b53ed68-c29a-4201-99dd-3bd8ca64f195 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bfae3521-0115-4d23-939f-5ef65fa177af · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Advise: Symbolism and external knowledge for decoding advertisements
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c227290e-e616-4373-acf9-30ac7471b18c · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 292164a1-f20b-40fb-be92-cb7a07e7acf3 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models BERTScore: Evaluating Text Generation with BERT
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65f006a9-25be-4729-ba27-caed35ba339d · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Towards automatic learning of procedures from web instructional videos
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b76d694a-719c-482f-9df9-111cba72f2d2 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models End-to-end dense video captioning with masked transformer
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4bbfa2c7-ae93-4542-a974-454ca2e2bf22 · outbound
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models Unresolved cited work
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70d83a25-4c97-4471-8c44-fd8ede4213ab · inbound
SERUM: State Extraction and Refinement for User Modeling Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.