Pith. sign in

Paper Citation Record · LEDGER

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs

As of 11 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2501.05884.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05884 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:11:43.604791Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:55:47.952638Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:53:16.138644Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ef59cbb-19cf-4fa8-8a33-afce7dd858ff · outbound

This paper cites GPT-4 Technical Report.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:41.564124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:41.564124Z digest=sha256:8c49958e65d3b56c941e179e7f512f2b7eab854d41ca8653e7f09940daa25d36

Observation e8d9590f-a5dc-4eb1-b9ab-db618f84aee6 · outbound

This paper cites Automatic compo- sition techniques for video production.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Automatic compo- sition techniques for video production

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:50.944837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:41.604460Z digest=sha256:113afd7f2a4888ba19c74874c9c8c16271df0f163128aebfd4b26440e1d6acfb

Observation 7bafc4f3-2248-4f1c-bf2d-0422e962dbd6 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:41.644751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:41.644751Z digest=sha256:89e79b72892034e9848bc136b656e9ca32b85d16a39b1cf68889287d8dc02b99

Observation da08f5a4-d4f5-4e88-a3bb-450529e29c6a · outbound

This paper cites Automatic editing of footage from multi- ple social cameras.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Automatic editing of footage from multi- ple social cameras

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:50.705147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:41.685512Z digest=sha256:db20b785c1c167d48d065fd101389e3be47aa1aa4fa41e884dde8c3622d3153f

Observation c599bf2b-25a6-4bb0-8ae0-d1d5527a9d3b · outbound

This paper cites The anatomy of video editing: A dataset and benchmark suite for ai-assisted video editing.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs The anatomy of video editing: A dataset and benchmark suite for ai-assisted video editing

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:50.555100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:41.734983Z digest=sha256:5a40f616d833cd336626ed15302a55569371bda3f115bb5ba3bed65a027bba3a

Observation c255b114-0ebe-4de1-9032-4414eeee0ead · outbound

This paper cites Qwen Technical Report.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Qwen Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:41.784756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:41.784756Z digest=sha256:e714b2bddd5db421a3c19aad6c58c641c6f1d1b80e938326819839bd3e50cbaa

Observation d759664b-1851-4df2-8d40-987f3a4e7c5d · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:41.820830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:41.820830Z digest=sha256:ec46c235db47650e15c591546fa4398d7a20a72dc1eda0df97355a6cd23e091d

Observation 5f443ae9-1d0c-4250-9b85-1b1b3bf6ce4b · outbound

This paper cites InternLM2 Technical Report.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs InternLM2 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:41.874830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:41.874830Z digest=sha256:47312a48070debeddd18c0a82842debcfc67cf081c6e808c46f08bb2739c482f

Observation 089054f9-eded-455b-b4cb-89d509ba8427 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:50.314749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:41.924754Z digest=sha256:351d505e320f96d8149c65bbea83b1945b9a9c6daf45d43950c579f37b1cd833

Observation 208d3cbf-e6d5-468c-9f20-451b2971b825 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:41.974749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:41.974749Z digest=sha256:de27caa77c4da9c524f5bfca5c894a9f2bb1a52f7fd7a7cd0a2ef203af75e318

Observation 853e0aa9-3949-4a08-8d69-08c27e9260e6 · outbound

This paper cites A video retrieval and se- quencing system.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs A video retrieval and se- quencing system

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:50.031370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:42.019886Z digest=sha256:bb44950fff29178566d921d3c14f747df8f423442a42a9e97326d47660529c38

Observation 06a03372-1428-46c4-b7c5-4d2a7b54d5bb · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.064391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.064391Z digest=sha256:92f963a7c996d8c2758544e72518c168782e5152d61b294e90ca6f5141a657da

Observation 5746b209-7f32-421f-994d-b229879251c0 · outbound

This paper cites Slowfast networks for video recognition.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Slowfast networks for video recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.118283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.118283Z digest=sha256:35f8d8dee202fd03783f1c5bc78dbc796b985050140d9589f4a5347e388ae44d

Observation 4c9ee0e4-a239-41b4-8a4d-371a50d25083 · outbound

This paper cites Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.144841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.144841Z digest=sha256:cd3fd6726dcbd695dfb525e6a10bde6cadd6c594710a9e69cdb34175dcd9b6ec

Observation 4a44f93b-4ac3-4fb0-a08d-52468befb4f3 · outbound

This paper cites LITA: Language Instructed Temporal-Localization Assistant.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs LITA: Language Instructed Temporal-Localization Assistant

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.184866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.184866Z digest=sha256:4b809e579859869de4d0f24da6f6c8bd3c8a881569c1b7dfa6e8c33c79594f74

Observation 7d94592c-7175-4fa4-801d-a5db8850843c · outbound

This paper cites B-script: Transcript-based b- roll video editing with recommendations.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs B-script: Transcript-based b- roll video editing with recommendations

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:49.764759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:42.224757Z digest=sha256:1125f8e9686bb0a1b6ef14e256f279573fc084fa001e7c5eb7aa861de98a307a

Observation d890f49d-2825-480d-b7c6-70095f624f50 · outbound

This paper cites Mistral 7B.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Mistral 7B

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.266610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.266610Z digest=sha256:454a072dfe16a2e0c6b798672b38350d5a71768ca7d317a261253c7c732ab5d3

Observation 059148e2-6031-4b0c-bc83-244e0f7e2fcd · outbound

This paper cites Chunkyedit: Text-first video interview editing via chunking.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Chunkyedit: Text-first video interview editing via chunking

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:49.576364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:42.312566Z digest=sha256:e05e5710f8ee886deeb3ce87b3a41ea6b62fd03aaa1edfe67227f3664559e85b

Observation 8cd4bcc0-63d0-40eb-bad9-8cac88c9a874 · outbound

This paper cites Computational video editing for dialogue-driven scenes.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Computational video editing for dialogue-driven scenes

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:49.424839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:42.335237Z digest=sha256:6d8cc1d0340d3870e6ec9f6ffea1f1a7ea311d9e762eb1d6ac5ba5b282753522

Observation 6eeb1c1e-3b03-41d6-9be3-71d31f15a380 · outbound

This paper cites Llava-next: Tack- ling multi-image, video, and 3d in large multimodal models,.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Llava-next: Tack- ling multi-image, video, and 3d in large multimodal models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:49.294759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:42.375405Z digest=sha256:da44e9b375fb258136b0a2f072c4d2bfd34e4a52653fa2b9fd478b7cfe9609e3

Observation 498b4893-3921-4e28-880d-c47ec4f430be · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.433458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.433458Z digest=sha256:c5e494d8ccd95fd63bed8243a9c5caad6c2d7dc00bcce0f064468b91b7ea75f8

Observation 1b58b57d-cdaa-47b7-969d-fcfee4a0f4d9 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.474493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.474493Z digest=sha256:4ea53985684c8ec8dabdf011f2fadf3d18b99cecd0596f6ed53189ff99affddc

Observation 727e52ea-efb0-4812-98a9-6999c0d20fdc · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Rouge: A package for automatic evaluation of summaries

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.524855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.524855Z digest=sha256:61f104954e260328d830e6509e140f8aaa21afde915f46bd841194b58c8ac02e

Observation b898a39e-13c7-4cec-abab-0f50348c5fa9 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:48.895633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:42.537068Z digest=sha256:9c1a733aeab2250c1d9e7c7f11ce623ed73161f50f9c3920072fea22bab6b34f

Observation 96de7618-1e4f-4548-9661-cf37a7f3cb4a · outbound

This paper cites Visual instruction tuning.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Visual instruction tuning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:48.734740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:42.541051Z digest=sha256:badb270bdd7dcd5097c490529e910b415c0b4e1ad00eae80ad0a0875fb80b6b8

Observation 15d7e393-627f-4ddb-a0f1-1164b426fc0d · outbound

This paper cites Decoupled Weight Decay Regularization.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Decoupled Weight Decay Regularization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.554764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.554764Z digest=sha256:e44e025808a9f355c6537c39a59d5153ff3a8fb30613ee9bed75e2970f18d605

Observation d96706d9-c741-4f5f-8162-e5c890f6e44f · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.604754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.604754Z digest=sha256:70a4ccc7f29fdcc61e432b1dc777871bec60d8594a4ba5eeb7b25498742187f1

Observation 8893fbe4-de43-400e-b3e2-bd7511f87085 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.647673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.647673Z digest=sha256:66d9a837e62310401027c4b9f49043f15fdc78e32727a0698b32734540eb6059

Observation 89758d2e-baae-449b-bb4b-b99d2ff77202 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.674840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.674840Z digest=sha256:38409f5f52cd30e6ca0ff8a61fc5aee4d3b45f028498c9e3e762ea7fb773eccd

Observation 677cd09f-62e4-4d93-86f2-881f51f640b7 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.724755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.724755Z digest=sha256:42147c329bc97fb69edb5832146c8f62eab30c07fbc06132aad2ed6760e78f83

Observation 6c28810d-18a2-40b6-8324-33d8461d93ff · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Bleu: a method for automatic evaluation of machine translation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.774756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.774756Z digest=sha256:c24eb73154d9503e73b3b2b210fad5eefabf34059ac15a3b7d60fea09ff78ba6

Observation 46364018-9714-48f9-bb8a-28214a1ccd92 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Learning transferable visual models from natural language supervi- sion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.834757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.834757Z digest=sha256:ecd4ea7cfffa1878a65a99cc23b16c2443ff9e4ba684a63d209ac963e2b139d4

Observation e64acbcb-4953-4405-8109-4ffaea00db0a · outbound

This paper cites Emscore: Evaluating video captioning via coarse-grained and fine-grained embed- ding matching.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Emscore: Evaluating video captioning via coarse-grained and fine-grained embed- ding matching

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:48.384744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:42.884757Z digest=sha256:3f26140330375a24a605148ca96d9be42a1a5a71d68ff2688ae129350973dd7b

Observation 6b269b59-b155-4c60-94a6-5c7bfa6043ee · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Moviechat: From dense token to sparse memory for long video understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.934760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.934760Z digest=sha256:24effc36f468a66460f7acceaf2c0b5649c2eec132d645e80a7bbbea7cc27187

Observation 0cb35494-cb89-4e33-8cc0-c2d8dee9ebf0 · outbound

This paper cites Transnet v2: An effective deep network architecture for fast shot transition detection.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Transnet v2: An effective deep network architecture for fast shot transition detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:48.164757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:42.984864Z digest=sha256:50685b1555717c110322632e1a085c02512015faa57b0bdb4b0954aa8ac1fc67

Observation 3617100a-a5dd-4470-8bd8-4e514af46c5f · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:43.034832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:43.034832Z digest=sha256:2f0c71263bd14926e4c490bfd8daf6bc0493481426f06516865e0871f372892f

Observation 843fe2f6-4de6-4808-9930-896b05c6cfa4 · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities, 2023.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Internlm: A multilingual language model with progressively enhanced capabilities, 2023

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:48.004750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:43.079982Z digest=sha256:cb2ec1688df7b022d492e6fc4d70c5e8f31a9b526502f06aa498f404485ac9d9

Observation bd0a5bbe-5dd6-47b6-ade7-a4c6a767b80e · outbound

This paper cites Quickcut: An interactive tool for editing nar- rated video.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Quickcut: An interactive tool for editing nar- rated video

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:47.834754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:43.124835Z digest=sha256:4d0cac878361d1cd644478abec2b00f16f21771659c8f24afec90a8e26af4907

Observation 7995648b-aca2-4542-98fa-09e1d9eb5c0e · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Cider: Consensus-based image description evalua- tion

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:43.154751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:43.154751Z digest=sha256:a38071c912199b44217cb2bdc512d319f32a514a95cb02bec553b024579a8c0f

Observation 64cf7ca5-e0cb-4965-8221-1f69004aba3c · outbound

This paper cites Write-a-video: computational video montage from themed text.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Write-a-video: computational video montage from themed text

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:47.544830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:43.198033Z digest=sha256:48baa5843886026082fb00558f32cc4cc5d76add9cca7d7386318ee71a790b9d

Observation 237f9a83-513c-4f34-8631-9ee49d10f275 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:43.248927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:43.248927Z digest=sha256:16e41d3c5b3db310b53a7315661d00dcb367e27ccaa62b28272a7cadd4dd0a8c

Observation 37b00b6f-ed4b-4f12-a71b-0bd0db4171f7 · outbound

This paper cites Transcript to video: Efficient clip sequencing from texts.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Transcript to video: Efficient clip sequencing from texts

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:47.394754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:43.274760Z digest=sha256:4919962b1c6b54b13ef80098c73f247537b9d14db0799b9a00a5afe8f4055c51

Observation eed9aaa0-a14e-4ab8-910a-49e95b0dda84 · outbound

This paper cites Qwen2 Technical Report.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Qwen2 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:43.310157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:43.310157Z digest=sha256:1602af4a70899f8b8d375849d5f53ae40a81c213362c01cb5e23b0598d4817fa

Observation dee47068-bf73-4a73-9bd7-7a65d99457a6 · outbound

This paper cites Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:43.330137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:43.330137Z digest=sha256:b11432bbf3fbf57d05a033be189785269daddc8c12beb766cb45457c5b636feb

Observation 45a1ae2e-43c5-431f-93d3-48045b7b4696 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:43.373939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:43.373939Z digest=sha256:ae8682c451dccd98eb72159512198f527e4ade36ea8a6ba2735bdfea9bf19694

Observation 45bb5982-8e7d-4c9b-a061-3bc67d406622 · outbound

This paper cites Sigmoid loss for language image pre-training.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Sigmoid loss for language image pre-training

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:47.234754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:43.417003Z digest=sha256:3a2df971759e8555e547d0e428eff091d477f3f6985211db3ba2feb048663b44

Observation ba32da1a-03e9-47a8-ade7-c1f6c7f396d5 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:43.461799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:43.461799Z digest=sha256:fd6b6150e3ca27913006cba668c5880b085877faa9da28fd7c53a82369c3c7e7

Observation e4b8ac7c-b76f-4cd1-aaea-ca08e083fe13 · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:43.509358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:43.509358Z digest=sha256:cc51a4e48142aa8e22e38b8adee3974aae5618825ecf57d9351e687cfb3493bd

Observation c46444b1-690a-4b9e-b4e8-d731a12a85a8 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Llava- next: A strong zero-shot video understanding model, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:47.094857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:43.544792Z digest=sha256:3ab761a22c3ade5dc5dbb51c4cdbe2379ac1239ce2823fe2c00d31ee7230a347

Observation 26b32438-c7da-4ca9-99ad-25ce653d2e15 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:43.570791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:43.570791Z digest=sha256:277996dadffbd0f8833333f5e7adc79bd5a0a08575dcca703420b4d2146f8831

Observation 96c5407a-953f-4924-8b49-ea07a4cc202f · outbound

This paper cites voice_over_track.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs voice_over_track

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:11:46.934921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:11:43.604791Z digest=sha256:d873710420ce9aaccb5a38a8aead428d7edc69b5b5597207df8f9f093a164f05

Pith citing papers

Observation 918d5ffd-be6e-4cff-b8e5-889155676255 · inbound

Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation cites this paper.

Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:47.952638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:55:47.952638Z digest=sha256:6cef97c410a6afaff11a5ec79fda59304fcb24390334e184406f5bbd4d6e35ec

Observation 3bf8dae2-f173-4cda-a356-df17af850545 · inbound

MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production cites this paper.

MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:13:29.970104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T09:10:31.245928Z digest=sha256:2ba32ae761747320e561253ead705578b368a12b0b4915feb372ad0f1d2b3b5a

Observation 9e84cef2-21f2-4502-9311-cbef42c81319 · inbound

KGEdit: Ambiguity-Aware Knowledge Graphs for Training-Free Precise Video Generation and Editing cites this paper.

KGEdit: Ambiguity-Aware Knowledge Graphs for Training-Free Precise Video Generation and Editing Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:16.140099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T08:46:56.390567Z digest=sha256:bddfb8dfeb8f2c9b5ceeb2276b8d941a64dbd9257bf29e1a7cc6eee712c4b56e