Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:11:43.604791Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2501.05884.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:11:43.604791Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:55:47.952638Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T08:53:16.138644Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1ef59cbb-19cf-4fa8-8a33-afce7dd858ff · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8d9590f-a5dc-4eb1-b9ab-db618f84aee6 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Automatic compo- sition techniques for video production
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7bafc4f3-2248-4f1c-bf2d-0422e962dbd6 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Flamingo: a visual language model for few-shot learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da08f5a4-d4f5-4e88-a3bb-450529e29c6a · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Automatic editing of footage from multi- ple social cameras
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c599bf2b-25a6-4bb0-8ae0-d1d5527a9d3b · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs The anatomy of video editing: A dataset and benchmark suite for ai-assisted video editing
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c255b114-0ebe-4de1-9032-4414eeee0ead · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Qwen Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d759664b-1851-4df2-8d40-987f3a4e7c5d · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f443ae9-1d0c-4250-9b85-1b1b3bf6ce4b · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs InternLM2 Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 089054f9-eded-455b-b4cb-89d509ba8427 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 208d3cbf-e6d5-468c-9f20-451b2971b825 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 853e0aa9-3949-4a08-8d69-08c27e9260e6 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs A video retrieval and se- quencing system
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 06a03372-1428-46c4-b7c5-4d2a7b54d5bb · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5746b209-7f32-421f-994d-b229879251c0 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Slowfast networks for video recognition
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c9ee0e4-a239-41b4-8a4d-371a50d25083 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a44f93b-4ac3-4fb0-a08d-52468befb4f3 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs LITA: Language Instructed Temporal-Localization Assistant
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d94592c-7175-4fa4-801d-a5db8850843c · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs B-script: Transcript-based b- roll video editing with recommendations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d890f49d-2825-480d-b7c6-70095f624f50 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Mistral 7B
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 059148e2-6031-4b0c-bc83-244e0f7e2fcd · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Chunkyedit: Text-first video interview editing via chunking
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8cd4bcc0-63d0-40eb-bad9-8cac88c9a874 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Computational video editing for dialogue-driven scenes
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6eeb1c1e-3b03-41d6-9be3-71d31f15a380 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Llava-next: Tack- ling multi-image, video, and 3d in large multimodal models,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 498b4893-3921-4e28-880d-c47ec4f430be · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b58b57d-cdaa-47b7-969d-fcfee4a0f4d9 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 727e52ea-efb0-4812-98a9-6999c0d20fdc · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Rouge: A package for automatic evaluation of summaries
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b898a39e-13c7-4cec-abab-0f50348c5fa9 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 96de7618-1e4f-4548-9661-cf37a7f3cb4a · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Visual instruction tuning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 15d7e393-627f-4ddb-a0f1-1164b426fc0d · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Decoupled Weight Decay Regularization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96706d9-c741-4f5f-8162-e5c890f6e44f · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs SGDR: Stochastic Gradient Descent with Warm Restarts
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8893fbe4-de43-400e-b3e2-bd7511f87085 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89758d2e-baae-449b-bb4b-b99d2ff77202 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 677cd09f-62e4-4d93-86f2-881f51f640b7 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c28810d-18a2-40b6-8324-33d8461d93ff · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Bleu: a method for automatic evaluation of machine translation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46364018-9714-48f9-bb8a-28214a1ccd92 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Learning transferable visual models from natural language supervi- sion
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e64acbcb-4953-4405-8109-4ffaea00db0a · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Emscore: Evaluating video captioning via coarse-grained and fine-grained embed- ding matching
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6b269b59-b155-4c60-94a6-5c7bfa6043ee · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Moviechat: From dense token to sparse memory for long video understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cb35494-cb89-4e33-8cc0-c2d8dee9ebf0 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Transnet v2: An effective deep network architecture for fast shot transition detection
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3617100a-a5dd-4470-8bd8-4e514af46c5f · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 843fe2f6-4de6-4808-9930-896b05c6cfa4 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Internlm: A multilingual language model with progressively enhanced capabilities, 2023
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bd0a5bbe-5dd6-47b6-ade7-a4c6a767b80e · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Quickcut: An interactive tool for editing nar- rated video
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7995648b-aca2-4542-98fa-09e1d9eb5c0e · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Cider: Consensus-based image description evalua- tion
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64cf7ca5-e0cb-4965-8221-1f69004aba3c · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Write-a-video: computational video montage from themed text
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 237f9a83-513c-4f34-8631-9ee49d10f275 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37b00b6f-ed4b-4f12-a71b-0bd0db4171f7 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Transcript to video: Efficient clip sequencing from texts
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eed9aaa0-a14e-4ab8-910a-49e95b0dda84 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Qwen2 Technical Report
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dee47068-bf73-4a73-9bd7-7a65d99457a6 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45a1ae2e-43c5-431f-93d3-48045b7b4696 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45bb5982-8e7d-4c9b-a061-3bc67d406622 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Sigmoid loss for language image pre-training
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ba32da1a-03e9-47a8-ade7-c1f6c7f396d5 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4b8ac7c-b76f-4cd1-aaea-ca08e083fe13 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c46444b1-690a-4b9e-b4e8-d731a12a85a8 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Llava- next: A strong zero-shot video understanding model, 2024
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 26b32438-c7da-4ca9-99ad-25ce653d2e15 · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96c5407a-953f-4924-8b49-ea07a4cc202f · outbound
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs voice_over_track
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 918d5ffd-be6e-4cff-b8e5-889155676255 · inbound
Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bf8dae2-f173-4cda-a356-df17af850545 · inbound
MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9e84cef2-21f2-4502-9311-cbef42c81319 · inbound
KGEdit: Ambiguity-Aware Knowledge Graphs for Training-Free Precise Video Generation and Editing Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.