Pith. sign in

REVIEW 3 major objections 6 minor 91 references

SimTube: Generating Simulated Video Comments through Multimodal AI and User Personas

T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read SimTube generates simulated audience comments before a video is published, and its comments are rated more relevant, believable, and helpful than actual popular comments.

desk verdict A solid system paper whose usefulness survives the weak baseline, but the abstract's comparative claim needs to be walked back or re-tested on a fair sample of real comments. read the letter →

arxiv 2411.09577 v2 pith:RAKNLSI7 submitted 2024-11-14 cs.HC

classification cs.HC
keywords simulatedvideocommentsaudiencefeedbackmultimodalunderstandinguserpersonaslargelanguagemodelscommentgenerationcontentcreation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces SimTube, a generative system that produces simulated audience comments for videos that have not yet been published. Its goal is to let creators obtain diverse, believable feedback before release, instead of waiting for real viewers to comment. The authors' central claim is that SimTube's comments are not only relevant, believable, and diverse but often more detailed and informative than actual audience comments. If that holds, creators could use pre-publication simulated feedback to refine their current video, and novice creators with small audiences would gain a scalable source of input.

What carries the argument

The load-bearing mechanism is a three-stage backend pipeline. First, Video Understanding transcribes the audio with Whisper, captions one-second segments by giving a vision-language model (LLaVA-NeXT 13B) a four-frame 'panel' paired with the dialogue, and then summarizes the captions and transcript into a video summary and keywords using a long-context LLM (Claude 1.6). Second, Persona Query embeds the keywords and over 8,000 PersonaChat personas, ranks them by cosine similarity, and keeps the top 30 relevant personas. Third, Comment Generation prompts an LLM with few-shot examples of real video comments, the video metadata, summary, keywords, and personas to produce primary comments and thread replies, and it also generates comments from user-defined personas and replies in threads. The UI lets creators expand threads and craft their own personas.

What would settle it

Take the same eight videos and ask blind raters to evaluate generated comments against real comments sampled from the full comment distribution, or matched for length, and check whether generated comments still score higher on helpfulness. A second check is to ask experienced creators to identify which comments were AI-generated; if they can reliably distinguish them, the believability claim weakens.

Watch

Extended reading notes

Core claim

On the paper's own terms, SimTube shows that multimodal video understanding plus persona-based generation can produce comments that match or exceed real video comments in relevance, believability, and helpfulness. In a crowd-sourced study with 25 workers rating comments on eight videos, both the full system and a no-persona ablated version scored significantly higher than real comments on relevance and believability, and the full system also scored significantly higher on helpfulness; the difference between the full and no-persona versions was not statistically significant. Automatic metrics showed generated comments have higher word-level diversity and stronger word- and semantic-level relevance to the video, while real comments retain greater semantic diversity. The paper frames this as evidence that simulated comments can supply preliminary inspiration and actionable feedback before publication, complementing rather than replacing real human comments.

Load-bearing premise

The evaluation assumes that real comments sampled from the top 1000 popular comments on each video fairly represent genuine audience feedback; the paper itself notes these comments are often brief and offer limited informative value for creators, so the comparison may exaggerate SimTube's advantage.

Editorial extensions

If this is right

  • Creators can gather audience-style feedback before publishing, potentially improving the current video rather than only future episodes.
  • Novice creators with few real viewers gain a scalable source of simulated input that does not depend on posting first.
  • Persona crafting and thread expansion let creators ask what their target audience would think and follow up with specific questions.
  • Because the no-persona version performed nearly as well as the full system on the headline metrics, simpler pipelines may already provide useful feedback.
  • Simulated comments are typically more constructive and less harsh than real comments, which may reduce emotional harm during review.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the same pipeline could be applied to thumbnails, titles, or rough-cut segments to predict audience reactions beyond comments.
  • Testable extension: comparing generated comments against a length-matched or full-distribution sample of real comments would show whether the helpfulness advantage persists when the baseline is not dominated by brief, laudatory popular comments.
  • Potential risk: if creators rely heavily on simulated feedback, they may overfit to a generic, polite audience and miss the unconventional reactions that drive real engagement; the paper itself recommends simulated comments as a supplement, not a replacement.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper presents SimTube, a full-stack system that generates simulated YouTube comments from videos by combining multimodal video understanding (Whisper transcription, LLaVA-NeXT frame captioning, LLM summarization) with persona-based prompting. The user interface lets creators upload videos, inspect generated comments, expand threads, and craft custom personas. The evaluation includes a crowd-sourced study (25 Upwork raters) comparing relevance, believability, and helpfulness of SimTube-generated comments against actual YouTube comments; automatic diversity and relevance metrics (Distinct N-gram, Self-BLEU, BERTScore, ROUGE, LLM Eval); and a qualitative user study with eight experienced creators. The paper claims that SimTube's generated comments are not only relevant, believable, and diverse but often more detailed and informative than actual audience comments.

Significance. If the central claim holds, SimTube offers video creators a scalable way to obtain audience-like feedback before publication, which could benefit content iteration and reduce reliance on post-publication comments. The paper's strengths include a complete implemented system, a multi-part evaluation that includes independent human crowd ratings, and a qualitative user study with experienced creators that surfaces concrete workflow integration scenarios. The crowd-sourced ratings are an external grounding for the relevance and believability of the generated comments. However, the headline comparative claim depends on a real-comment baseline that the authors themselves characterize as brief and laudatory, and the automatic relevance metrics are partly circular because the generated comments are produced from the same video summary used as the evaluation reference. These issues must be resolved before the comparative claim is fully established.

major comments (3)
  1. [Sec. 6.1.1 and Sec. 6.1.3] The real-comment baseline is not a neutral sample of audience feedback: Section 6.1.1 states that comments were randomly sampled from the top 1000 popular comments per video, ranked by YouTube meta-ratings, and Section 6.1.3 admits that these popular comments 'are often brief and offer limited informative value for creators' and are 'sometimes purely laudatory.' Because the headline claim in the abstract is explicitly comparative ('more detailed and informative than actual audience comments'), this curated baseline biases the comparison in favor of SimTube. The crowd-sourced study should be repeated with a random sample of all comments (after filtering), or with comments matched for length and informativeness, to support the claimed advantage.
  2. [Sec. 6.2.2 and Sec. 5.1.3] The automatic relevance metrics (ROUGE, BERTScore, LLM Eval) compare generated comments against the video summary and keywords that the pipeline itself generates in Section 5.1.3. Since the comment generation prompt in Section 5.3.1 is given this same summary and keywords, high word-level and semantic overlap with that summary is partly by construction, making the reported relevance advantage over real comments unsurprising. The independent crowd-sourced relevance ratings provide stronger evidence, but the automatic metrics should be reframed as measuring consistency with the system's own summary rather than as independent evidence of relevance.
  3. [Sec. 5.2.1 and Abstract] The abstract claims that user personas are 'derived from a broad and diverse corpus of audience demographics,' but the actual persona source is PersonaChat, a general persona dataset from a dialogue task, not a corpus of YouTube audience demographics. Furthermore, Section 6.1.3 reports that the full system with personas did not significantly outperform the No-persona ablation, and Section 7.3.2 reports mixed user reactions to persona-based comments. The paper's contribution claim for the persona module is therefore not supported by the presented evidence; either the wording should be corrected or additional analyses (e.g., measuring persona diversity or demonstrating a significant benefit) should be provided.
minor comments (6)
  1. [Sec. 6.1.1] The sentence 'For each video, we sampled 30 comments from three categories' is ambiguous: it is not clear whether 30 comments are sampled per category (90 total per video) or 30 comments total across all three categories. This should be clarified because it determines the number of ratings collected.
  2. [Sec. 6.2.2] The statement that 'ground-truth video summaries were compiled from 30 User Study participants using Claude 3.5 Sonnet' is unclear: it should specify how the participant-provided summaries were aggregated or processed by the LLM, and whether the same model was used in the LLM evaluation, which could introduce a self-preference bias.
  3. [Sec. 5.1.2] The claim that the frame-panel approach is 'a unique approach' is an overstatement; paneling multiple frames for VLM input has been explored in prior video understanding work. I recommend softening this wording.
  4. [Sec. 5.2.2] There is a typo: 'calcualte' should be 'calculate.' Similar typos appear in Appendix C ('isscooping') and Section 4.3 ('comment' missing a noun).
  5. [Figure 5] The description of the Self-BLEU plot in Section 6.2.1 is confusing: the text refers to 'the right side of the plot titled Self-BLEU' and then mentions 'the left half' without clear labels in the figure. Please make the figure and caption self-explanatory.
  6. [Sec. 5.3] Section 5.3 states the system directly generates 30 comments with 70% primary and 30% thread comments, but Section 6.1.1 also mentions 30 comments per video for evaluation; please clarify whether these are the same 30 comments and how they relate to the 671 and 1141 comments used in the automatic metrics.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the central claim rests on independent crowd-sourced human ratings and qualitative studies, with only a minor non-load-bearing self-citation.

full rationale

SimTube is a system paper with no formal derivation chain to reduce. The headline comparative claim that generated comments are 'more detailed and informative than actual audience comments' is primarily supported by the crowd-sourced study (Sec. 6.1), in which 25 Upwork raters independently rated Relevance, Believability, and Helpfulness on a 7-point Likert scale, and by qualitative interviews with experienced creators (Sec. 7); these are external human judgments, not quantities derived from the system's own inputs. The automatic metrics in Sec. 6.2 are supplementary rather than load-bearing. Although generated comments are produced from an LLM-generated video summary (Sec. 5.1.3, 5.3.1) and relevance is measured against ground-truth video summaries compiled from user-study participants via Claude 3.5 Sonnet (Sec. 6.2.2), the reference summaries are not identical to the generation input, the LLM evaluators are cross-model (Claude 3.5 Sonnet and GPT-4), and the authors explicitly cite work on LLM self-preference bias [2] in mitigation; this is a validity caveat, not a by-construction reduction. The one self-citation, [42] (LLM-Eval, co-authored by Yen-Ting Lin), supports the choice of LLM evaluation but is not load-bearing because the human ratings independently establish the main result. The paper itself flags a baseline limitation in Sec. 6.1.3: the real-comment baseline drawn from the top-1000 popular comments is 'often brief and offer limited informative value for creators,' which is a generalizability caveat about the strength of the comparative claim, not a circularity. Overall, no significant circularity is present.

Assumptions & free parameters 5 free parameters · 6 assumptions · 0 invented entities

The system introduces no invented physical or conceptual entities; it reuses existing models and datasets. The central claim depends on several hand-chosen hyperparameters (persona count, comment count, frame panel size, Whisper model, context window) and on assumptions about the reliability of pretrained models, the representativeness of PersonaChat, and the fairness of the popular-comments baseline. Free parameters are design choices, not fitted values, but they shape output quality and are not derived from first principles.

free parameters (5)
  • top_k_personas_per_video = 30
    Section 5.2.2 selects the top 30 personas by cosine similarity between keyword embeddings and persona embeddings; this hand-chosen number directly controls the diversity and relevance of generated comments.
  • generated_comments_per_video = 30 (70% primary, 30% thread)
    Section 5.3 sets the initial batch to 30 comments split 70/30 between primary and thread comments, a design choice not derived from data.
  • frame_panel_size = 4 consecutive frames
    Section 5.1.2 assembles four consecutive frames into a 'panel' input for the VLM; this choice affects temporal context captured in captions.
  • whisper_model_size = medium
    Section 5.1.1 selects Whisper-medium 'for its balanced efficiency and performance'; a hand-chosen trade-off that affects transcription quality.
  • llm_context_window_requirement = 200K tokens
    Section 5.1.3 relies on Claude's 200K-token context to fit all captions and transcripts; the choice of model and context limit determines how much video content is summarized.
assumptions (6)
  • domain assumption Whisper-medium, LLaVA-NeXT 13B, Claude (200K context), and text-embedding-3-small produce sufficiently accurate video summaries, captions, transcripts, and relevance scores for comment generation to be useful.
    Sections 5.1 and 5.2 rely entirely on these pretrained models; the paper does not measure their individual error rates or failure modes.
  • domain assumption PersonaChat personas, ranked by cosine similarity to video keywords, represent plausible YouTube commenter demographics and backgrounds.
    Section 5.2 uses PersonaChat as the persona corpus; Appendix F shows that some retrieved personas produce confusing, off-topic comments, indicating this assumption can fail.
  • domain assumption Crowd workers on Upwork can judge relevance, believability, and helpfulness of comments consistently after watching the video and passing a quiz.
    Section 6.1.2 describes the protocol but does not report inter-rater reliability or per-question variance.
  • ad hoc to paper The top 1000 popular YouTube comments per video, after filtering non-English and offensive content, are a fair baseline representing 'real audience comments.'
    Section 6.1.1; the authors acknowledge these popular comments are often brief and laudatory (Section 6.1.3), which biases the comparison in favor of SimTube.
  • domain assumption ROUGE, BERTScore, and LLM-based relevance scores are valid proxies for comment relevance.
    Section 6.2 applies these automatic metrics to summarize relevance; the text relies on prior validation literature but does not validate them on this comment task.
  • standard math Wilcoxon signed-rank test with Bonferroni correction is appropriate for the matched crowd ratings.
    Section 6.1.3 uses this non-parametric paired test, which is standard for ordinal Likert-scale data.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SimTube: Generating Simulated Video Comments through Multimodal AI and User Personas." pith.science (2026). https://pith.science/paper/RAKNLSI7

@misc{pith2026241109577,
  author       = {Pith},
  title        = {Pith review of: SimTube: Generating Simulated Video Comments through Multimodal AI and User Personas},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RAKNLSI7}},
  note         = {Machine review of arXiv:2411.09577}
}
read the original abstract

Audience feedback is crucial for refining video content, yet it typically comes after publication, limiting creators' ability to make timely adjustments. To bridge this gap, we introduce SimTube, a generative AI system designed to simulate audience feedback in the form of video comments before a video's release. SimTube features a computational pipeline that integrates multimodal data from the video-such as visuals, audio, and metadata-with user personas derived from a broad and diverse corpus of audience demographics, generating varied and contextually relevant feedback. Furthermore, the system's UI allows creators to explore and customize the simulated comments. Through a comprehensive evaluation-comprising quantitative analysis, crowd-sourced assessments, and qualitative user studies-we show that SimTube's generated comments are not only relevant, believable, and diverse but often more detailed and informative than actual audience comments, highlighting its potential to help creators refine their content before release.

Figures

Figures reproduced from arXiv: 2411.09577 by the authors.

Figure 1
Figure 1. SimTube (A) generates synthetic content-based video comments and (B) supports user-steerable comment generation from users’ responding or defined persona, providing preliminary inspiration and insight for iterating video. Abstract Audience feedback is crucial for refining video content, yet it typ￾ically comes after publication, limiting creators’ ability to make timely adjustments. To bridge this gap, we introduce … view at source ↗
Figure 2
Figure 2. Users upload a video via the Upload Video Page, and (A) view the comments generated by SimTube on the Simulated [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Thread Expansion: Upon receipt of (A) the user’s reply, The thread is expanded by both (B) the user’s reply and (C) the generated response of the commenter. Persona Crafting: (D) Upon user specification of a persona, (E) a new comment is generated in alignment with the video content and the user-defined persona. approach that considers the temporal aspects of video while align￾ing visual and audio tracks. First, we … view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: SimTube leverage VLM and Whisper to process videos in visual and audio components per modality. Personas are [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: This plot includes quantitative results from Crowd￾Sourced Study and other automatic metrics measuring diver￾sity such as BERTScore, Distinct N-grams, and Self-BLEU. 6.2.1 Diversity A diverse set of comments offers varied perspectives, encompassing different word choic…
Figure 6
Figure 6. Figure 6: This figure contains quantitative results from au [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 8
Figure 8. Figure 8: The frames are captured in a cooking YouTube [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 7
Figure 7. Figure 7: The frames are captured in a gaming YouTube [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

91 extracted references · 38 canonical work pages

  1. [1]

    Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision . 2425–2433

  2. [2]

    Bowman and Shi Feng

    Arjun Panickssery and Samuel R. Bowman and Shi Feng. 2024. LLM Evaluators Recognize and Favor Their Own Generations. arXiv:2404.13076 [cs.CL] https: //arxiv.org/abs/2404.13076

  3. [3]

    Andrei Barbu, Alexander Bridge, Zachary Burchill, Dan Coroian, Sven Dickinson, Sanja Fidler, Aaron Michaux, Sam Mussman, Siddharth Narayanaswamy, Dhaval Salvi, Lara Schmidt, Jiangnan Shangguan, Jeffrey Mark Siskind, Jarrell Waggoner, Song Wang, Jinlian Wei, Yifan Yin, and Zhiqi Zhang. 2012. Video In Sentences Out. arXiv:1204.2742 [cs.CV] https://arxiv.org...

  4. [4]

    Kobus Barnard. 2016. Computational Methods for Integrating Vision and Language. Synthesis Lectures on Computer Vision 6 (04 2016), 1–227. https: //doi.org/10.2200/S00705ED1V01Y201602COV007

  5. [5]

    Karim Benharrak, Tim Zindulka, Florian Lehmann, Hendrik Heuer, and Daniel Buschek. 2024. Writer-Defined AI Personas for On-Demand Feedback Generation. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24, Vol. 11). ACM, 1–18. https://doi.org/10.1145/3613904.3642406

  6. [6]

    Simion-Vlad Bogolin, Ioana Croitoru, and Marius Leordeanu. 2020. A hierar- chical approach to vision-based language generation: from simple sentences to complex natural language. In Proceedings of the 28th International Conference on Computational Linguistics, Donia Scott, Nuria Bel, and Chengqing Zong (Eds.). In- ternational Committee on Computational Li...

  7. [7]

    Tom B Brown. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165 (2020)

  8. [8]

    Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. 2023. Sparks of Artificial General Intelligence: Early experiments with GPT-4. arXiv:2303.12712 [cs.CL] https://arxiv.org/abs/2303.12712

Show all 91 references
  1. [9]

    My Way of Telling a Story

    Khyathi Chandu, Shrimai Prabhumoye, Ruslan Salakhutdinov, and Alan W Black. 2019. “My Way of Telling a Story”: Persona based Grounded Story Generation. In Proceedings of the Second Workshop on Storytelling , Francis Fer- raro, Ting-Hao ‘Kenneth’ Huang, Stephanie M. Lukin, and ...

  2. [10]

    Sorina Chelaru, Claudia Orellana-Rodriguez, and Ismail Sengor Altingovde. 2014. How useful is social feedback for learning to rank YouTube videos? World Wide Web 17, 5 (2014), 997–1025. https://doi.org/10.1007/s11280-013-0258-9

  3. [11]

    Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, Li Yuan, Yu Qiao, Dahua Lin, Feng Zhao, and Jiaqi Wang. 2024. ShareGPT4Video: Improving Video Understanding and Generation with Better Captions. arXiv:2406.043...

  4. [12]

    Ahmed Kharrufa Colin Dodds. 2024. Show-and-Tell: An Interface for Delivering Rich Feedback upon Creative Media Artefacts. Multimodal Technol. Interact (2024). https://www.mdpi.com/2414-4088/8/3/23

  5. [13]

    Digital Marketing Institute. 2024. 14 Ways to Grow Your YouTube Channel. https://digitalmarketinginstitute.com/. https://digitalmarketinginstitute.com/ blog/10-ways-to-grow-your-youtube-channel-in-2018 Accessed: 2024-10-01

  6. [14]

    Pierre Dognin, Igor Melnyk, Youssef Mroueh, Inkit Padhi, Mattia Rigotti, Jarret Ross, Yair Schiff, Richard A Young, and Brian Belgodere. 2022. Image captioning as an assistive technology: Lessons learned from vizwiz 2020 challenge. Journal of Artificial Intelligence Research 7...

  7. [15]

    Chaoqun Duan, Lei Cui, Shuming Ma, Furu Wei, Conghui Zhu, and Tiejun Zhao. 2020. Multimodal Matching Transformer for Live Commenting. arXiv:2002.02649 [cs.CL] https://arxiv.org/abs/2002.02649

  8. [16]

    Ilana Dubovi and Iris Tabak. 2020. An empirical analysis of knowledge co- construction in YouTube comments. Computers & Education 156 (2020), 103939. https://doi.org/10.1016/j.compedu.2020.103939

  9. [17]

    Ali Farhadi, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth. 2010. Every Picture Tells a Story: Generating Sentences from Images. In Computer Vision – ECCV 2010 , Kostas Daniilidis, Petros Maragos, and Nikos Paragios ...

  10. [18]

    Shansan Gong, Mukai Li, Jiangtao Feng, Zhiyong Wu, and LingPeng Kong. 2022. Diffuseq: Sequence to sequence text generation with diffusion models. arXiv preprint arXiv:2210.08933 (2022)

  11. [19]

    Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022. News Summarization and Evaluation in the Era of GPT-3. ArXiv abs/2209.12356 (2022). https://api. semanticscholar.org/CorpusID:252532176

  12. [20]

    Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh

  13. [21]

    Sergio Guadarrama, Niveda Krishnamoorthy, Girish Malkarnenkar, Subhashini Venugopalan, Raymond Mooney, Trevor Darrell, and Kate Saenko. 2013. YouTube2Text: Recognizing and Describing Arbitrary Activities Using Semantic Hierarchies and Zero-Shot Recognition. In 2013 IEEE Intern...

  14. [22]

    Danna Gurari, Yinan Zhao, Meng Zhang, and Nilavra Bhattacharya. 2020. Cap- tioning images taken by people who are blind. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XVII

  15. [23]

    Perttu Hämäläinen, Mikke Tavast, and Anton Kunnari. 2023. Evaluating Large Language Models in Generating Synthetic HCI Research Data: a Case Study. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Com...

  16. [24]

    John Hattie and Helen Timperley. 2007. The Power of Feedback. Review of Edu- cational Research 77, 1 (2007), 81–112. https://doi.org/10.3102/003465430298487 arXiv:https://doi.org/10.3102/003465430298487

  17. [25]

    Maria Holmbom. 2015. The YouTuber: A Qualitative Study of Popular Content Creators. Dissertation. Umeå University. https://urn.kb.se/resolve?urn=urn:nbn: se:umu:diva-105388

  18. [26]

    Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin, and Hamid Laga

    MD. Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin, and Hamid Laga

  19. [27]

    De-An Huang, Vignesh Ramanathan, Dhruv Mahajan, Lorenzo Torresani, Manohar Paluri, Li Fei-Fei, and Juan Carlos Niebles. 2018. What Makes a Video a Video: Analyzing Temporal Information in Video Understanding Models and Datasets. In Proceedings of the IEEE Conference on Compute...

  20. [28]

    Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, and Margaret Mitchell

    Ting-Hao Kenneth Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, C. Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, and Margaret Mitchell. 2016. Visual Storytell...

  21. [29]

    Mina Huh, Saelyne Yang, and Yi-Hao Peng. [n.d.]. Xiang’Anthony’Chen, Young- Ho Kim, and Amy Pavel. 2023. AVscript: Accessible Video Editing with Audio- Visual Scripts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–17

  22. [30]

    Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara

  23. [31]

    josepmartins. [n.d.]. boring-avatar. https://github.com/boringdesigners/boring- avatars

  24. [32]

    Taehyeong Kim, Min-Oh Heo, Seonil Son, Kyoung-Wha Park, and Byoung-Tak Zhang. 2019. GLAC Net: GLocal Attention Cascading Networks for Multi-image Cued Story Generation. arXiv:1805.10973 [cs.CL] https://arxiv.org/abs/1805. 10973

  25. [33]

    Avraham N Kluger and Angelo DeNisi. 1996. The effects of feedback interventions on performance: a historical review, a meta-analysis, and a preliminary feedback intervention theory. Psychological bulletin 119, 2 (1996), 254

  26. [34]

    Saydulu Kolasani. 2023. Optimizing Natural Language Processing, Large Lan- guage Models (LLMs) for Efficient Customer Service, and hyper-personalization to enable sustainable growth and revenue. Transactions on Latest Trends in Artifi- cial Intelligence 4, 4 (2023). https://ij...

  27. [35]

    Kulkarni, Michael S

    Chinmay E. Kulkarni, Michael S. Bernstein, and Scott R. Klemmer. 2015. PeerStu- dio: Rapid Peer Feedback Emphasizes Revision and Improves Performance. In Proceedings of the Second (2015) ACM Conference on Learning @ Scale (Vancouver, BC, Canada) (L@S ’15). Association for Comp...

  28. [36]

    Girish Kulkarni, Visruth Premraj, Sagnik Dhar, Siming Li, Yejin Choi, Alexander C Berg, and Tamara L Berg. 2011. Baby talk: Understanding and generating simple image descriptions. In CVPR 2011. 1601–1608. https://doi.org/10.1109/CVPR.2011. 5995466

  29. [37]

    Polina Kuznetsova, Vicente Ordonez, Alexander Berg, Tamara Berg, and Yejin Choi. 2012. Collective Generation of Natural Image Descriptions. InProceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Haizhou Li, Chin-Yew L...

  30. [38]

    Mina Lee, Percy Liang, and Qian Yang. 2022. CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model Capabilities. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI ’22). Association for...

  31. [39]

    Lee and Amy J

    Michael J. Lee and Amy J. Ko. 2011. Personifying programming tool feedback improves novice programmers’ learning. In Proceedings of the Seventh Interna- tional Workshop on Computing Education Research (Providence, Rhode Island, USA) (ICER ’11). Association for Computing Machin...

  32. [40]

    Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016. A Diversity-Promoting Objective Function for Neural Conversation Models. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lan...

  33. [41]

    Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74–81

  34. [42]

    Yen-Ting Lin and Yun-Nung Chen. 2023. LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models. In Proceedings of the 5th Workshop on NLP for Conversational AI (NLP4ConvAI 2023), Yun-Nung Chen and Abhinav Rastogi (Eds.)....

  35. [43]

    Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634 (2023)

  36. [44]

    Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to explain: Multi- modal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems 35 ...

  37. [45]

    Shuming Ma, Lei Cui, Damai Dai, Furu Wei, and Xu Sun. 2018. LiveBot: Generating Live Video Comments Based on Visual and Textual Contexts. arXiv:1809.04938 [cs.CL] https://arxiv.org/abs/1809.04938

  38. [46]

    Ali Malik, Mike Wu, Vrinda Vasavada, Jinpeng Song, Madison Coots, John Mitchell, Noah Goodman, and Chris Piech. 2021. Generative Grading: Near Human-level Accuracy for Automated Feedback on Richly Structured Problems. arXiv:1905.09916 [cs.LG] https://arxiv.org/abs/1905.09916

  39. [47]

    Meta. 2022. PersonaChat. https://www.kaggle.com/datasets/atharvjairath/ personachat Accessed: 2024-10-01

  40. [48]

    Mathewson, Jaylen Pittman, and Richard Evans

    Piotr Mirowski, Kory W. Mathewson, Jaylen Pittman, and Richard Evans. 2023. Co-Writing Screenplays and Theatre Scripts with Language Models: Evaluation by Industry Professionals. In Proceedings of the 2023 CHI Conference on Human 14 Factors in Computing Systems (Hamburg, Germa...

  41. [49]

    Margaret Mitchell, Xufeng Han, Jesse Dodge, Alyssa Mensch, Amit Goyal, Alex Berg, Kota Yamaguchi, Tamara Berg, Karl Stratos, and Hal Daumé. 2012. Midge: generating image descriptions from computer vision detections. In Proceedings of the 13th Conference of the European Chapter...

  42. [50]

    Hamed Nilforoshan and Eugene Wu. 2018. Leveraging Quality Prediction Models for Automatic Writing Feedback.Proceedings of the International AAAI Conference on Web and Social Media 12, 1 (Jun. 2018). https://doi.org/10.1609/icwsm.v12i1. 14998

  43. [51]

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Bal- com, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jef...

  44. [52]

    Jim Owens. 2023. Video Production Handbook (7th ed.). Routledge. https: //doi.org/10.4324/9781003251323

  45. [53]

    Keivalya Pandya and Mehfuza Holia. 2023. Automating Customer Service us- ing LangChain: Building custom open-source GPT Chatbot for organizations. arXiv:2310.05421 [cs.CL] https://arxiv.org/abs/2310.05421

  46. [54]

    O’Brien, Carrie J

    Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative Agents: Interactive Simulacra of Human Behavior. arXiv:2304.03442 [cs.HC]

  47. [55]

    Bernstein

    Joon Sung Park, Lindsay Popowski, Carrie Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2022. Social Simulacra: Creating Populated Prototypes for Social Computing Systems. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Techn...

  48. [56]

    Goldman, Björn Hartmann, and Maneesh Agrawala

    Amy Pavel, Dan B. Goldman, Björn Hartmann, and Maneesh Agrawala. 2016. VidCrit: Video-based Asynchronous Video Review. In Proceedings of the 29th Annual Symposium on User Interface Software and Technology (Tokyo, Japan) (UIST ’16). Association for Computing Machinery, New York...

  49. [57]

    Ludovica Piro, Tommaso Bianchi, Luca Alessandrelli, Andrea Chizzola, Daniela Casiraghi, Susanna Sancassani, and Nicola Gatti. 2024. MyLearningTalk: An LLM-Based Intelligent Tutoring System. In Web Engineering, Kostas Stefanidis, Kari Systä, Maristella Matera, Sebastian Heil, H...

  50. [58]

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2022. Robust Speech Recognition via Large-Scale Weak Supervision. arXiv:2212.04356 [eess.AS]

  51. [59]

    Gonzalo Ramos and Ravin Balakrishnan. 2003. Fluid interaction techniques for the control and annotation of digital video. In Proceedings of the 16th Annual ACM Symposium on User Interface Software and Technology (Vancouver, Canada) (UIST ’03). Association for Computing Machine...

  52. [60]

    Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. 2023. Character-LLM: A Trainable Agent for Role-Playing. arXiv:2310.10158 [cs.CL] https://arxiv.org/ abs/2310.10158

  53. [61]

    Zhiqiang Shen, Jianguo Li, Zhou Su, Minjun Li, Yurong Chen, Yu-Gang Jiang, and Xiangyang Xue. 2017. Weakly supervised dense video captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 1916–1924

  54. [62]

    Stefan Siersdorfer, Sergiu Chelaru, Wolfgang Nejdl, and Jose San Pedro. 2010. How useful are your comments? analyzing and predicting youtube comments and comment ratings. In Proceedings of the 19th International Conference on World Wide Web (Raleigh, North Carolina, USA) (WWW ...

  55. [63]

    Marko Smilevski, Ilija Lalkovski, and Gjorgji Madjarov. 2018. Stories for Images- in-Sequence by Using Visual and Narrative Components . Springer International Publishing, 148–159. https://doi.org/10.1007/978-3-030-00825-3_13

  56. [64]

    Jingkuan Song, Yuyu Guo, Lianli Gao, Xuelong Li, Alan Hanjalic, and Heng Tao Shen. 2018. From deterministic to generative: Multimodal stochastic RNNs for video captioning. IEEE transactions on neural networks and learning systems 30, 10 (2018), 3047–3058

  57. [65]

    SSA. 2023. USA Baby Name Dataset. https://www.ssa.gov/OACT/babynames/ limits.html

  58. [66]

    Marie Stevenson and Aek Phakiti. 2014. The effects of computer-generated feedback on the quality of writing. Assessing Writing 19 (2014), 51–65. https: //doi.org/10.1016/j.asw.2013.11.007 Feedback in Writing: Issues and Challenges

  59. [67]

    Chun Chet Tan, Yu-Gang Jiang, and Chong-Wah Ngo. 2011. Towards textually describing complex video contents with audio-visual concept classifiers. In Pro- ceedings of the 19th ACM International Conference on Multimedia (Scottsdale, Arizona, USA) (MM ’11). Association for Comput...

  60. [68]

    Mingkang Tang, Zhanyu Wang, Zhenhua LIU, Fengyun Rao, Dian Li, and Xiu Li. 2021. CLIP4Caption: CLIP for Video Caption. In Proceedings of the 29th ACM International Conference on Multimedia (Virtual Event, China) (MM ’21). Association for Computing Machinery, New York, NY, USA,...

  61. [69]

    Mike Thelwall, Pardeep Sud, and Farida Vis. 2012. Commenting on YouTube videos: From Guatemalan rock to el big bang. Journal of the American society for information science and technology 63, 3 (2012), 616–629

  62. [70]

    Thematic Analysis Inc. 2024. Thematic Comment Analysis. https://getthematic. com/product/comment-analyzer/ Accessed: 2024-10-01

  63. [71]

    Yu-Min Tseng, Yu-Chao Huang, Teng-Yun Hsiao, Wei-Lin Chen, Chao-Wei Huang, Yu Meng, and Yun-Nung Chen. 2024. Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization. arXiv:2406.01171 [cs.CL] https://arxiv.org/ abs/2406.01171

  64. [72]

    Tess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C Derry, Mina Huh, and Amy Pavel. 2024. Making Short-Form Videos Accessible with Hierarchical Video Sum- maries. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Assoc...

  65. [73]

    VEED. 2024. VEED. https://www.veed.io/ Accessed: 2024-10-01

  66. [75]

    Bryan Wang, Yuliang Li, Zhaoyang Lv, Haijun Xia, Yan Xu, and Raj Sodhi. 2024. LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video Editing. arXiv:2402.10294 [cs.HC] https://arxiv.org/abs/2402.10294

  67. [76]

    Yaqing Wang, Quanming Yao, James T Kwok, and Lionel M Ni. 2020. Generalizing from a few examples: A survey on few-shot learning. ACM computing surveys (csur) 53, 3 (2020), 1–34

  68. [77]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824–24837

  69. [78]

    Hao Wu, Gareth J. F. Jones, and Francois Pitie. 2020. Response to Live- Bot: Generating Live Video Comments Based on Visual and Textual Contexts. arXiv:2006.03022 [cs.CL] https://arxiv.org/abs/2006.03022

  70. [79]

    Michihiro Yasunaga, Jure Leskovec, and Percy Liang. 2022. Linkbert: Pretraining language models with document links. arXiv preprint arXiv:2203.15827 (2022)

  71. [80]

    Dongwook Yoon, Nicholas Chen, François Guimbretière, and Abigail Sellen

  72. [81]

    YouTube. 2023. Made on YouTube. https://blog.youtube/news-and-events/made- on-youtube-2023/

  73. [82]

    Kuo-Hao Zeng, Tseng-Hung Chen, Ching-Yao Chuang, Yuan-Hong Liao, Juan Car- los Niebles, and Min Sun. 2017. Leveraging Video Descriptions to Learn Video Question Answering. Proceedings of the AAAI Conference on Artificial Intelligence 31, 1 (Feb. 2017). https://doi.org/10.1609/...

  74. [83]

    Zehua Zeng, Neng Gao, Cong Xue, and Chenyang Tu. 2021. PLVCG: A Pretraining Based Model for Live Video Comment Generation. In Advances in Knowledge Dis- covery and Data Mining , Kamal Karlapalem, Hong Cheng, Naren Ramakrishnan, R. K. Agrawal, P. Krishna Reddy, Jaideep Srivasta...

  75. [84]

    Shu Zhang, Xinyi Yang, Yihao Feng, Can Qin, Chia-Chih Chen, Ning Yu, Zeyuan Chen, Huan Wang, Silvio Savarese, Stefano Ermon, Caiming Xiong, and Ran Xu

  76. [85]

    Weinberger, and Yoav Artzi

    Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi

  77. [86]

    Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018. Texygen: A benchmarking platform for text generation models. In The 41st international ACM SIGIR conference on research & development in information retrieval. 1097–1100. 16

  78. [89]

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

    HIVE: Harnessing Human Feedback for Instructional Visual Editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 9026–9036

  79. [2014]

    In Proceedings of the 27th Annual ACM Symposium on User Interface Software and Technology(Honolulu, Hawaii, USA)(UIST ’14)

    RichReview: blending ink, speech, and gesture to support collaborative document review. In Proceedings of the 27th Annual ACM Symposium on User Interface Software and Technology(Honolulu, Hawaii, USA)(UIST ’14). Association for Computing Machinery, New York, NY, USA, 481–490. ...

  80. [2017]

    In Proceedings of the IEEE conference on computer vision and pattern recognition

    Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition . 6904–6913

  81. [2019]

    ACM Comput

    A Comprehensive Survey of Deep Learning for Image Captioning. ACM Comput. Surv. 51, 6, Article 118 (feb 2019), 36 pages. https://doi.org/10.1145/ 3295748

  82. [2020]

    InInternational Confer- ence on Learning Representations

    BERTScore: Evaluating Text Generation with BERT. InInternational Confer- ence on Learning Representations. https://openreview.net/forum?id=SkeHuCVFDr

  83. [2024]

    arXiv:2305.02547 [cs.CL] https://arxiv.org/abs/2305.02547

    PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits. arXiv:2305.02547 [cs.CL] https://arxiv.org/abs/2305.02547

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.