REVIEW 4 major objections 6 minor 1 cited by
Lotus: Creating Short Videos From Long Videos With Abstractive and Extractive Summarization
T0 review · 4 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Lotus claims that combining AI-written narration with original footage lets creators make short videos from long ones in minutes, with results they rate comparable to their usual tools.
desk verdict A solid HCI systems paper that genuinely combines abstractive and extractive video summarization; the headline preference result isn't significant and the matching pipeline needs validation, but the work deserves review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the clip-transcript matching and blending pipeline. For matching, each long-form clip is scored against a visual concept from the short-form script using four signals: speech similarity via nomic-embed-text embeddings, keyframe similarity via CLIP, a GPT-4o perceptual score that rates roughly 25 candidate clips from their keyframes and speech, and position-based alignment that rewards clips whose relative position in the long video matches the concept's relative position in the script. For blending, abstractive and extractive segments are compared on speech similarity, the CLIP visual connection between a segment's speech and its keyframe, noun-phrase coverage, and position; when the best extractive match is close, the system generates permutations and picks the one with the highest coherence as measured by GPT-3 loss. This machinery turns an abstractive draft into a mixed video and gives the editing interface concrete alternatives to offer the user.
What would settle it
Take a set of long videos, run Lotus's automatic clip-transcript matching, and have independent human annotators judge each assigned clip against its transcript sentence; if a substantial fraction of assignments are judged mismatched or the human agreement with GPT-4o's scores is low, the automatic draft's quality claim is falsified. A second test: run the full user study with a 'broken draft' condition in which the initial script is paired with randomly chosen clips, and measure whether satisfaction and final quality remain comparable.
Extended reading notes
Core claim
Lotus's central claim is that an editing system can productively combine abstractive and extractive summarization in one workflow, and that creators will prefer working this way over starting from scratch. The pipeline first generates a short-form transcript from the long-form transcript with GPT-4o, extracts visual concepts, matches long-form clips to those concepts with a weighted scoring function, and synthesizes narration with ElevenLabs to produce an initial abstractive draft. It then scores abstractive segments against extractive segments and, where the difference is small, lets the user choose between versions so the final video can blend newly written narration with original footage. The user study found all eight participants wanted to use Lotus in the future, found the generated draft a useful starting point, and rated their Lotus-made videos comparable in quality to those from their existing tools; the results evaluation found the mixed method preferred, though not significantly, with preferences varying by video type.
Load-bearing premise
The load-bearing premise is that GPT-4o's perceptual scoring of which long-form clips match the written narration is good enough that the initial abstractive draft is genuinely usable; if those judgments are unreliable, the draft would be mismatched and the positive user results would partly reflect participants' willingness to repair a weak starting point.
Editorial extensions
If this is right
- Creators can produce a usable short-form draft in far less than the hours-to-days manual editors in the formative study reported.
- The mixed approach is not uniformly best: for videos with a prominent on-camera narrator, extractive or mixed clips win, while for narrator-light footage abstractive summaries are preferred.
- The initial abstractive draft serves as a scaffold: participants kept many of its clips and used alignment and search to add original footage around it.
- Switching a clip between abstractive and extractive modes gives creators a way to fix narration they find unnatural without losing the visual.
- If the preference for the mixed method holds beyond the study's sample, short-form tools should offer the blend as a default rather than forcing one summarization strategy.
Reading between the lines
- The positive user-study results have not been separated from the quality of the automatic clip-transcript matching; a controlled study in which the initial draft is replaced by a random or mismatched draft would show how much of the value comes from the draft versus the interface.
- The genre-dependence found in the ranking study suggests an adaptive default: the system could decide per video whether to start abstractive, extractive, or mixed.
- Because participants wanted a horizontal timeline and finer trimming, the approach could be packaged as a module inside existing editors rather than as a standalone tool.
- The reliance on a synthetic voice was the most-cited weakness; improvements in expressive speech generation would likely shift more creators toward abstractive clips.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Lotus is a system that helps creators repurpose long-form videos into short-form videos by combining abstractive summarization (generated narration with matched visuals) and extractive summarization (original clips). The paper reports a formative study with professional editors, a two-stage generation pipeline (short-form transcript generation, visual concept extraction, clip-transcript matching with a weighted scoring function, and blending of abstractive and extractive clips), and an editing interface. Evaluation consists of a results evaluation (12 raters, 6 videos, three conditions: abstractive, extractive, and mixed) and a within-subjects user study (8 participants) comparing Lotus to participants' existing editing tools. The paper claims that the mixed method was more preferred than either pure method (though not significantly), that all participants would use Lotus in the future, and that participants created videos comparable in quality to those made with existing tools.
Significance. The paper addresses a genuinely important and under-served problem: the time-consuming work of repurposing long videos into short-form social media content. The proposed combination of abstractive and extractive summarization is a reasonable and timely design idea, and the system is described in unusual detail, including full prompts in the appendix, the use of publicly available models, and a clear interface description. If the effectiveness claims survive revision, Lotus would be a valuable contribution to the video-authoring and intelligent-user-interfaces literatures. The main strengths are the formative study that grounds the design goals, the explicit separation of abstractive and extractive capabilities with user control, and the open reporting of the pipeline. The main weaknesses are the lack of validation of the automatic clip-transcript matching and the over-interpretation of non-significant quantitative results.
major comments (4)
- [§4.3.3 and Table 1] The clip-transcript matching is a core load-bearing component that supports design goal G1 and the claim that the initial draft reduces creator effort, but the paper provides no evidence that GPT-4o's scoring of candidate clips agrees with human judgments of relevance. The user study interaction log in Table 1 shows very high numbers of trim operations (e.g., P7: 157 trims, P1: 65 trims), which is consistent with substantial manual repair of the initial draft. Without a measure of how often the initially matched clips were retained or replaced, the reader cannot tell whether the positive user reactions reflect the quality of the automatic draft or the participants' tolerance for fixing a weak draft. Please add either a human-rated validation of the matching component (e.g., relevance ratings for matched versus alternative clips) or a quantitative retention/replacement analysis from the logged user interactions.
- [§5, Figure 5] The headline result of the results evaluation, that the mixed method is 'more preferred' than the abstractive or extractive methods, is not statistically significant: with 12 raters, the mean ranks are 1.88, 2.06, and 2.04. The abstract and introduction state that the results evaluation 'demonstrates the benefit of flexibility,' but with no inferential test and no effect size, this is an overstatement. Please report a proper analysis (e.g., Friedman test with post-hoc comparisons, or a Bayesian equivalent), and adjust the wording of the conclusions to describe the pattern as suggestive and exploratory rather than demonstrative.
- [§6.3, Figures 6 and 7] The user study makes comparative claims, such as participants created videos of comparable quality 'without experiencing increased mental demand' and 'liked the editing process within Lotus over their existing tools,' but no statistical tests are reported for the TLX, Likert, or CSI ratings. With n=8, these claims are currently supported only by descriptive statistics and error bars. Please report paired non-parametric tests (e.g., Wilcoxon signed-rank) for the key comparisons, or explicitly limit those claims to qualitative observations from the interviews.
- [§4.3.4] The blending of abstractive and extractive segments depends on an 'empirically determined' threshold and a 'coherence score, determined by GPT-3 [26] loss,' but the paper does not describe the threshold value, the procedure used to set it, or how the GPT-3 loss is computed and applied. Because the mixed method is the central contribution, the reader cannot reproduce or independently assess the blending step. Please specify the threshold, the optimization or selection process, and the exact use of GPT-3 loss.
minor comments (6)
- [§4.2.1] The sentence 'This pane consists consists of the Original Video Player' contains a duplicated word; please fix.
- [§5, Table 2] The phrase 'Appendix Tableg 2' should be 'Appendix Table 2'.
- [§7] The first sentence of 'Integration With Existing Tools' repeats the same idea twice ('Lotus is not as fully functional as traditional video editing tools... does not provide a comprehensive suite of features'); please consolidate.
- [Throughout] The text contains many instances of 'fexibility' and 'efect' that appear to be spelling errors; please proofread to replace these with 'flexibility' and 'effect'.
- [§4.3.3] The formula for position-based alignment is typeset in a corrupted format ('????? ??? = 1 − ??? (visual concept)− ??? (long-form clip)'); the actual equation needs to be rendered correctly.
- [§9.1] The prompts in the appendix are a valuable contribution, but the 'GPT-4o Scoring' prompt would benefit from a short sentence explaining how the visual concept is embedded in the speech window, since the user message references an '[embedded visual concept]' that is not defined in the prompt text.
Circularity Check
No significant circularity: the system's evaluation rests on external user judgments and an independently re-implemented baseline, not on fitted values or self-citations.
full rationale
The paper's central claim is that a system combining abstractive and extractive summarization helps creators make short-form videos. This claim is supported by human evaluation: a results evaluation with 12 annotators comparing three methods, and a within-subjects user study (n=8) comparing Lotus against participants' existing editing tools. The methods' outputs are judged by external raters and participants, so the outcome is not derived from the system's own parameters. The clip-transcript matching in Section 4.3.3 uses a weighted score that includes GPT-4o perceptual ratings, and the blending threshold in Section 4.3.4 is described as 'empirically determined.' These are implementation parameters tuned by the authors, not predictions derived from the claim, and the paper does not present them as first-principles results. The unvalidated GPT-4o scoring could be a robustness or correctness concern, since noisy automatic matching might shift more work onto users, but that is an empirical weakness, not circularity. The paper's self-citations (AVscript [36], Video Digests [55], Rescribe [56], SceneSkim [54], and other works by Pavel et al.) appear in background and related-work contexts and are not load-bearing for the central evaluation; no uniqueness theorem or forced choice is imported from those citations. The extractive baseline is attributed to ROPE [68], whose authors are not among this paper's authors, and the paper states it re-implemented the method. Overall, no equation or claim reduces to its own inputs, and no fitted parameter is renamed as a prediction. The minor presence of the authors' own prior work in the background raises the circularity score only to 1.
Assumptions & free parameters
free parameters (2)
- Weighted scoring function weights =
not disclosed
- Blending threshold =
empirically determined, value not given
assumptions (4)
- domain assumption GPT-4o produces faithful summaries and reliable visual concept scores
- domain assumption CLIP embeddings and nomic embeddings capture semantic similarity between text and visuals
- domain assumption User self-reports in a small lab study predict real-world creative outcomes
- domain assumption The six YouTube trending videos are representative of long-form content for repurposing
Cite this review
Pith. "Pith review of Lotus: Creating Short Videos From Long Videos With Abstractive and Extractive Summarization." pith.science (2026). https://pith.science/paper/3Y6QQNU5
@misc{pith2026250207096,
author = {Pith},
title = {Pith review of: Lotus: Creating Short Videos From Long Videos With Abstractive and Extractive Summarization},
year = {2026},
howpublished = {\url{https://pith.science/paper/3Y6QQNU5}},
note = {Machine review of arXiv:2502.07096}
}
read the original abstract
Short-form videos are popular on platforms like TikTok and Instagram as they quickly capture viewers' attention. Many creators repurpose their long-form videos to produce short-form videos, but creators report that planning, extracting, and arranging clips from long-form videos is challenging. Currently, creators make extractive short-form videos composed of existing long-form video clips or abstractive short-form videos by adding newly recorded narration to visuals. While extractive videos maintain the original connection between audio and visuals, abstractive videos offer flexibility in selecting content to be included in a shorter time. We present Lotus, a system that combines both approaches to balance preserving the original content with flexibility over the content. Lotus first creates an abstractive short-form video by generating both a short-form script and its corresponding speech, then matching long-form video clips to the generated narration. Creators can then add extractive clips with an automated method or Lotus's editing interface. Lotus's interface can be used to further refine the short-form video. We compare short-form videos generated by Lotus with those using an extractive baseline method. In our user study, we compare creating short-form videos using Lotus to participants' existing practice.
Forward citations
Cited by 1 Pith paper
-
Transforming Podcast Preview Generation: From Expert Models to LLM-Based Systems
LLM-generated podcast previews beat a feature-engineered multi-model baseline in human evaluation and production A/B testing while processing episodes five times faster.
Reference graph
Works this paper leans on
-
[26]
Tom B Brown. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165 (2020)
arXiv 2020
-
[1]
2024. Adobe Premiere Pro. https://www.adobe.com/products/premiere.html
work page 2024
-
[2]
2024. CapCut. https://www.capcut.com/
work page 2024
- [3]
- [4]
- [5]
-
[6]
2024. FFmpeg. https://fmpeg.org/
work page 2024
- [7]
Show all 76 references
-
[8]
2024. GPT-4o. https://openai.com/index/hello-gpt-4o/
2024
-
[9]
2024. iMovie. https://support.apple.com/imovie
2024
-
[10]
Instagram Reels
2024. Instagram Reels. https://www.instagram.com/reels/
2024
-
[11]
2024. MoviePy. https://github.com/Zulko/moviepy
2024
-
[12]
OpusClip
2024. OpusClip. https://www.opus.pro/
2024
-
[13]
Resemble Enhance
2024. Resemble Enhance. https://github.com/resemble-ai/resemble-enhance
2024
-
[14]
2024. TikTok. https://www.tiktok.com/
2024
-
[15]
VEGAS Pro
2024. VEGAS Pro. https://www.vegascreativesoftware.com/us/vegas-pro/
2024
-
[16]
2024. Video 1. https://www.youtube.com/watch?v=EDap9qxb96k
2024
-
[17]
2024. Video 2. https://www.youtube.com/watch?v=IE-SyUTifag
2024
-
[18]
2024. Video 3. https://www.youtube.com/watch?v=m56-QeTP5cg
2024
-
[19]
2024. Video 4. https://www.youtube.com/watch?v=vyfJgJBB3Vk
2024
-
[20]
2024. Video 5. https://www.youtube.com/watch?v=JUfybRQc_1o
2024
-
[21]
2024. Video 6. https://www.youtube.com/watch?v=FSsRaoamxhM
2024
-
[22]
YouTube Short
2024. YouTube Short. https://www.youtube.com/shorts/
2024
-
[23]
Mehdi Allahyari, Seyedamin Pouriyeh, Mehdi Assef, Saeid Safaei, Elizabeth D Trippe, Juan B Gutierrez, and Krys Kochut. 2017. Text summarization techniques: a brief survey. arXiv preprint arXiv:1707.02268 (2017)
2017 arXiv
-
[24]
Dawit Mureja Argaw, Seunghyun Yoon, Fabian Caba Heilbron, Hanieh Deilam- salehy, Trung Bui, Zhaowen Wang, Franck Dernoncourt, and Joon Son Chung
-
[25]
Floraine Berthouzoz, Wilmot Li, and Maneesh Agrawala. 2012. Tools for placing cuts and transitions in interview video. ACM Transactions on Graphics (TOG) 31, 4 (2012), 1–8
2012
-
[27]
Juan Casares, A Chris Long, Brad A Myers, Rishi Bhatnagar, Scott M Stevens, Laura Dabbish, Dan Yocum, and Albert Corbett. 2002. Simplifying video editing using metadata. In Proceedings of the 4th conference on Designing interactive systems: processes, practices, methods, and t...
2002
-
[28]
Brandon Castellano. [n. d.]. PySceneDetect. https://github.com/Breakthrough/ PySceneDetect
-
[29]
Haozhe Chen, Run Chen, and Julia Hirschberg. 2024. EmoKnob: Enhance Voice Cloning with Fine-Grained Emotion Control. arXiv preprint arXiv:2410.00316 (2024)
2024 arXiv
-
[30]
Yoonseo Choi, Eun Jeong Kang, Seulgi Choi, Min Kyung Lee, and Juho Kim. 2024. Proxona: Leveraging LLM-Driven Personas to Enhance Creators’ Understanding of Their Audience. arXiv preprint arXiv:2408.10937 (2024)
2024 arXiv
-
[31]
Ohad Fried, Ayush Tewari, Michael Zollhöfer, Adam Finkelstein, Eli Shecht- man, Dan B Goldman, Kyle Genova, Zeyu Jin, Christian Theobalt, and Maneesh Agrawala. 2019. Text-based editing of talking-head video. ACM Transactions on IUI ’25, March 24–27, 2025, Cagliari, Italy Barua...
2019
-
[32]
Boqing Gong, Wei-Lun Chao, Kristen Grauman, and Fei Sha. 2014. Diverse sequential subset selection for supervised video summarization. Advances in neural information processing systems 27 (2014)
2014
-
[33]
Michael Gygli, Helmut Grabner, Hayko Riemenschneider, and Luc Van Gool
-
[34]
Matthew Honnibal, Ines Montani, Sofe Van Landeghem, and Adriane Boyd
-
[35]
Bernd Huber, Hijung Valentina Shin, Bryan Russell, Oliver Wang, and Gautham J Mysore. 2019. B-script: Transcript-based b-roll video editing with recommenda- tions. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 1–11
2019
-
[36]
Mina Huh, Saelyne Yang, Yi-Hao Peng, Xiang’Anthony’ Chen, Young-Ho Kim, and Amy Pavel. 2023. AVscript: Accessible Video Editing with Audio-Visual Scripts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–17
2023
-
[37]
Jef Hwang, Moto Hira, Caroline Chen, Xiaohui Zhang, Zhaoheng Ni, Guangzhi Sun, Pingchuan Ma, Ruizhe Huang, Vineel Pratap, Yuekai Zhang, Anurag Ku- mar, Chin-Yun Yu, Chuang Zhu, Chunxi Liu, Jacob Kahn, Mirco Ravanelli, Peng Sun, Shinji Watanabe, Yangyang Shi, Yumeng Tao, Robin ...
2023 arXiv
-
[38]
Hyunjoo Im, Billy Sung, Garim Lee, and Keegan Qi Xian Kok. 2023. Let voice assistants sound like a machine: Voice and task type efects on perceived fuency, competence, and consumer attitude. Computers in Human Behavior 145 (2023), 107791
2023
-
[39]
Shengpeng Ji, Jialong Zuo, Minghui Fang, Ziyue Jiang, Feiyang Chen, Xinyu Duan, Baoxing Huai, and Zhou Zhao. 2024. Textrolspeech: A text style control speech corpus with codec language text-to-speech models. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speec...
2024
-
[40]
Zhong Ji, Fang Jiao, Yanwei Pang, and Ling Shao. 2020. Deep attentive and semantic preserving video summarization. Neurocomputing 405 (2020), 200–207
2020
-
[41]
Zhong Ji, Kailin Xiong, Yanwei Pang, and Xuelong Li. 2019. Video summarization with attention-based encoder–decoder networks. IEEE Transactions on Circuits and Systems for Video Technology 30, 6 (2019), 1709–1717
2019
-
[42]
Haojian Jin, Yale Song, and Koji Yatani. 2017. Elasticplay: Interactive video summarization with dynamic time budgets. In Proceedings of the 25th ACM international conference on Multimedia. 1164–1172
2017
-
[43]
Zeyu Jin, Jia Jia, Qixin Wang, Kehan Li, Shuoyi Zhou, Songtao Zhou, Xiaoyu Qin, and Zhiyong Wu. 2024. Speechcraft: A fne-grained expressive speech dataset with natural language description. In Proceedings of the 32nd ACM International Conference on Multimedia. 1255–1264
2024
-
[44]
Juho Kim, Philip J Guo, Carrie J Cai, Shang-Wen Li, Krzysztof Z Gajos, and Robert C Miller. 2014. Data-driven interaction techniques for improving naviga- tion of educational videos. In Proceedings of the 27th annual ACM symposium on User interface software and technology. 563–572
2014
-
[45]
Mackenzie Leake, Abe Davis, Anh Truong, and Maneesh Agrawala. 2017. Com- putational video editing for dialogue-driven scenes. ACM Trans. Graph. 36, 4 (2017), 130–1
2017
-
[46]
Mackenzie Leake and Wilmot Li. 2024. ChunkyEdit: Text-frst video interview editing via chunking. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–16
2024
-
[47]
Mackenzie Leake, Hijung Valentina Shin, Joy O Kim, and Maneesh Agrawala. 2020. Generating audio-visual slideshows from text articles using word concreteness. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems. 1–11
2020
-
[48]
Daniel Li, Thomas Chen, Albert Tung, and Lydia B Chilton. 2021. Hierarchical summarization for longform spoken dialog. In The 34th Annual ACM Symposium on User Interface Software and Technology. 582–597
2021
-
[49]
Hui Lin and Vincent Ng. 2019. Abstractive summarization: A survey of the state of the art. In Proceedings of the AAAI conference on artifcial intelligence, Vol. 33. 9815–9822
2019
-
[50]
Rami Mubarak, Tariq Alsboui, Omar Alshaikh, Isa Inuwa-Dutse, Saad Khan, and Simon Parkinson. 2023. A survey on the detection and impacts of deepfakes in visual, audio, and textual formats. Ieee Access 11 (2023), 144497–144529
2023
-
[51]
Morris, Brandon Duderstadt, and Andriy Mulyar
Zach Nussbaum, John X. Morris, Brandon Duderstadt, and Andriy Mulyar. 2024. Nomic Embed: Training a Reproducible Long Context Text Embedder. Technical Report. arXiv:2402.01613 [cs.CL]
2024 arXiv
-
[52]
Mayu Otani, Yale Song, Yang Wang, et al. 2022. Video summarization overview. Foundations and Trends® in Computer Graphics and Vision 13, 4 (2022), 284–335
2022
-
[53]
Shruti Palaskar, Jindřich Libovický, Spandana Gella, and Florian Metze. 2019. Multimodal Abstractive Summarization for How2 Videos. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Anna Ko- rhonen, David Traum, and Lluís Màrquez (Eds....
2019 doi
-
[54]
Amy Pavel, Dan B Goldman, Björn Hartmann, and Maneesh Agrawala. 2015. Sceneskim: Searching and browsing movies using synchronized captions, scripts and plot summaries. In Proceedings of the 28th Annual ACM Symposium on User Interface Software & Technology. 181–190
2015
-
[55]
Amy Pavel, Colorado Reed, Björn Hartmann, and Maneesh Agrawala. 2014. Video digests: a browsable, skimmable format for informational lecture videos.. In UIST, Vol. 10. Citeseer, 2642918–2647400
2014
-
[56]
Amy Pavel, Gabriel Reyes, and Jefrey P Bigham. 2020. Rescribe: Authoring and automatically editing audio descriptions. In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology. 747–759
2020
-
[57]
Puyuan Peng, Po-Yao Huang, Shang-Wen Li, Abdelrahman Mohamed, and David Harwath. 2024. Voicecraft: Zero-shot speech editing and text-to-speech in the wild. arXiv preprint arXiv:2403.16973 (2024)
2024 arXiv
-
[58]
Alexis Plaquet and Hervé Bredin. 2023. Powerset multi-class cross entropy loss for neural speaker diarization. In Proc. INTERSPEECH 2023
2023
-
[59]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[60]
Steve Rubin, Floraine Berthouzoz, Gautham J Mysore, Wilmot Li, and Maneesh Agrawala. 2013. Content-based tools for editing audio stories. In Proceedings of the 26th annual ACM symposium on User interface software and technology. 113–122
2013
-
[61]
Shagan Sah, Sourabh Kulhare, Allison Gray, Subhashini Venugopalan, Emily Prud’Hommeaux, and Raymond Ptucha. 2017. Semantic text summarization of long videos. In 2017 IEEE Winter Conference on Applications of Computer Vision (W ACV). IEEE, 989–997
2017
-
[62]
Sarmiento, M
C. Sarmiento, M. Squicciarini, and J. Valdez Genao. 2024. Synthetic Content and its Implications for AI Policy: A Primer. Bernan Associates. https://books.google. com/books?id=Ym87EQAAQBAJ
2024
-
[63]
Hijung Valentina Shin, Wilmot Li, and Frédo Durand. 2016. Dynamic authoring of audio with linked scripts. In Proceedings of the 29th Annual Symposium on User Interface Software and Technology. 509–516
2016
-
[64]
Yale Song, Jordi Vallmitjana, Amanda Stent, and Alejandro Jaimes. 2015. Tvsum: Summarizing web videos using titles. In Proceedings of the IEEE conference on computer vision and pattern recognition. 5179–5187
2015
-
[65]
Bekzat Tilekbay, Saelyne Yang, Michal Adam Lewkowicz, Alex Suryapranata, and Juho Kim. 2024. ExpressEdit: Video Editing with Natural Language and Sketching. In Proceedings of the 29th International Conference on Intelligent User Interfaces. 515–536
2024
-
[66]
Anh Truong, Floraine Berthouzoz, Wilmot Li, and Maneesh Agrawala. 2016. Quickcut: An interactive tool for editing narrated video. In Proceedings of the 29th Annual Symposium on User Interface Software and Technology. 497–507
2016
-
[67]
Ba Tu Truong and Svetha Venkatesh. 2007. Video abstraction: A systematic review and classifcation. ACM transactions on multimedia computing, communications, and applications (TOMM) 3, 1 (2007), 3–es
2007
-
[68]
Bryan Wang, Zeyu Jin, and Gautham Mysore. 2022. Record Once, Post Every- where: Automatic Shortening of Audio Stories for Social Media. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. 1–11
2022
-
[69]
Bryan Wang, Yuliang Li, Zhaoyang Lv, Haijun Xia, Yan Xu, and Raj Sodhi. 2024. LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video Editing. In Proceedings of the 29th International Conference on Intelligent User Interfaces. 699–714
2024
-
[70]
Sitong Wang, Samia Menon, Tao Long, Keren Henderson, Dingzeyu Li, Kevin Crowston, Mark Hansen, Jefrey V Nickerson, and Lydia B Chilton. 2024. Reel- Framer: Human-AI co-creation for news-to-video translation. In Proceedings of the CHI Conference on Human Factors in Computing Sy...
2024
-
[71]
Sitong Wang, Zheng Ning, Anh Truong, Mira Dontcheva, Dingzeyu Li, and Lydia B Chilton. 2023. PodReels: Human-AI Co-Creation of Video Podcast Teasers. arXiv preprint arXiv:2311.05867 (2023)
2023 arXiv
-
[72]
Haijun Xia, Jennifer Jacobs, and Maneesh Agrawala. 2020. Crosscast: adding visuals to audio travel podcasts. In Proceedings of the 33rd annual ACM symposium on user interface software and technology. 735–746
2020
-
[73]
Here are the transcripts:
Bin Zhao, Xuelong Li, and Xiaoqiang Lu. 2017. Hierarchical recurrent neural network for video summarization. In Proceedings of the 25th ACM international conference on Multimedia. 863–871. Lotus: Creating Short Videos From Long Videos With Abstractive and Extractive Summarizat...
2017
-
[2014]
In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VII 13
Creating summaries from user videos. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VII 13. Springer, 505–520
2014
-
[2020]
spaCy: Industrial-strength Natural Language Processing in Python. (2020). https://doi.org/10.5281/zenodo.1212303
2020 doi
-
[2024]
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Scaling Up Video Summarization Pretraining with Large Language Mod- els. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 8332–8341
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.