{"id":"dff67442-dfc5-4704-aee2-07803e39d16b","arxiv_id":"2504.18805","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An agentic pipeline with iterative MLLM feedback composes scientific short-form videos from paper text and figures, beating a simple prompting baseline but remaining below human creator quality.","lead":"This paper introduces SciTalk, a multi-agent computer system that turns research papers into short, narrated videos by scripting, planning scenes, and adding text and effects. In tests on nine papers, it beat a simple one-prompt baseline on model-judged quality, but still produced videos judged below those made by human creators.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central iterative-refinement claim is scored by the same MLLM family that drives the loop; human ratings in Sec. 4.3 decline (SR 3.8→2.2, AVA 4.2→2.7), so the claimed improvement lacks validated evidence.","rationale":"I agree with the reader's identification of the weakest assumption: the human-alignment validity of the LLaVA-NeXT-Video feedback and evaluation agents. The concern is not mere disagreement with a model-based convention; it is a correctness risk because the optimizer and the reward signal are the same MLLM family. The paper's own Section 4.3 shows human ratings falling in SR and AVA while model ratings rise, which directly undermines the inference that iterative feedback improves perceived quality. I also note that the single-agent baseline in Figure 4 is model-only, so the comparative claim against simple prompting lacks human validation. Other issues—small human evaluation with two internal annotators, 4 of 9 seed papers from the authors' group, undefined exclusion of invalid outputs, and absence of statistical significance tests—are real but secondary; they weaken the empirical support without changing the core circularity. The framework itself is clearly described, the authors transparently acknowledge limitations, and the central claims are empirically testable, so CONDITIONAL remains the appropriate verdict pending external human evaluation and release of code and videos. My read does not move the reader's verdict.","tokens_in":17064,"tokens_out":4857,"duration_ms":51870,"concrete_test":"Run a blinded human evaluation with at least three external annotators, independent of the authors, on all 45 SciTalk videos plus the 9 single-agent baseline videos, using the Table 1 rubric with pre-registered inclusion/exclusion criteria. Compute inter-annotator agreement (e.g., Krippendorff's alpha), compare per-iteration human scores, and compare SciTalk iteration 1 against the baseline. If human scores do not rise across iterations, or if SciTalk does not beat the baseline on human-rated metrics, the abstract's claim that the iterative feedback loop and multi-agent pipeline improve scientific accuracy and engagement is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 states the Feedback Agents and the Evaluation Agent are 'each powered by a multi-modal LLM (MLLM)', and Section 4 identifies that model as LLaVA-NeXT-Video 34B. The Reflection Agent turns feedback scores into updated prompts for the next iteration, so the loop is optimizing against the same model family that also produces the evaluation scores. The primary quantitative evidence that iterations improve quality—Figure 3a feedback scores and Figure 4 model evaluation scores—therefore comes from an unvalidated, possibly self-referential proxy. Section 4.3 provides direct evidence against this proxy: while model scores in engagement and content accuracy improve or peak, human Scene Readability falls from 3.8 to 2.2 and Audio-Visual Alignment from 4.2 to 2.7 across iterations. The paper's own Discussion (Section 5) concedes that 'model-generated feedback often diverges from human evaluations'. Because the abstract and Section 4.2 claim that the iterative feedback loop and multi-agent system improve scientific accuracy and engagement, the load-bearing assumption is that LLaVA-NeXT-Video scores are faithful human-aligned judgments. That assumption is not established, and the provided human data contradict it. Additionally, Figure 4's single-agent baseline is described in the caption as a 'model-only average score', so the comparison against the simple prompting baseline also rests on the same unvalidated model judgments rather than human ratings.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SciTalk, a fully automatic multi-LLM agentic framework that generates short-form scientific videos from research papers. The pipeline decomposes the task into preprocessing, planning, editing, and feedback-evaluation stages, using specialized agents such as the Flashtalk Generator, Sceneplan Generator, and various editing assistants, and it composes final videos with a video editing library rather than generative video models. An iterative feedback loop, driven by LLaVA-NeXT-Video-based Feedback Agents and a Reflection Agent, refines agent prompts across five iterations. Experiments on nine papers report that SciTalk outperforms a single-agent prompting baseline in model-assessed content accuracy and engagement, though generated videos remain below human creator quality. The authors also present human evaluations and qualitative examples showing persistent issues such as text overlap and declining human-judged visual-audio synchronization over iterations.","tokens_in":17313,"tokens_out":4150,"duration_ms":42225,"significance":"If validated, SciTalk would be a useful contribution to automatic science dissemination, combining grounded source materials, specialized agent roles, and an explicitly designed iterative refinement loop. The use of a video composition library rather than generative models for the final assembly is a sensible design choice that avoids many visual artifacts. The paper also identifies a real and under-studied task, and the authors promise to release code, data, and generated videos, which would facilitate follow-up work. However, the central claims currently rest on an evaluation design that is self-referential and partially contradicted by the paper's own human data, so the significance of the reported results cannot yet be assessed as presented.","major_comments":[{"comment":"The central iterative-refinement claim is supported only by a self-referential evaluation loop. The Feedback Agents and the Evaluation Agent are each powered by LLaVA-NeXT-Video 34B, and the Reflection Agent revises prompts to increase the feedback scores from that same model family. The final videos are then rated by the same model family, so the upward trends in Figures 3a and 4 may reflect convergence to the evaluator's preferences rather than genuine quality improvement. Section 4.3 directly undermines the assumption that these model judgments are human-aligned: human Scene Readability falls from 3.8 to 2.2 and Audio-Visual Alignment from 4.2 to 2.7 across iterations while model scores rise. Consequently, the Abstract and Section 4.2 claims that the framework 'outperforms simple prompting methods' and 'improves compositional quality, coherence, and alignment with scientific messaging' are not supported by the current evidence.","section":"§4.3, Figure 4"},{"comment":"The multi-agent versus single-agent comparison is reported only as figure trends with no numeric values, no error bars for the single-agent baseline, and no significance tests. The caption identifies the baseline as a 'model-only average score,' but the main text never specifies how the single-agent prompt was constructed beyond saying it was 'simplified to produce direct end-to-end outputs,' nor does it report the distribution of scores or the number of videos included. Without quantitative reporting and statistical inference, the claim that the multi-agent system improves compositional quality, coherence, and scientific alignment is not established.","section":"§4.2, Figure 4"},{"comment":"The human evaluation uses only 'our two internal annotators' with no inter-annotator reliability statistics, no description of the annotation protocol, and no evidence of calibration against the rubric. Moreover, 'invalid outputs were excluded from aggregation' without a predefined exclusion criterion, and the sanity-check mechanism of Section 3.3 was deliberately disabled in the main experiments. This combination leaves open the possibility that exclusion was applied in a way that favors the framework. The paper should report the number and nature of excluded outputs, the exclusion rule, and the annotator agreement, e.g., Krippendorff's alpha or Cohen's kappa.","section":"§4 (Models & Implementation), §3.3"},{"comment":"The comparison against human creator videos uses different settings for the two evaluator types: human evaluations compare only the 1st iteration of SciTalk against creator videos, while model evaluations compare the average across all five iterations. This asymmetry makes the two panels of Figure 5 not directly comparable, and it further highlights the divergence between model and human judgments: the model evaluation favors SciTalk over creators on clarity and visual-audio synchronization, whereas the human evaluation favors creators. Because the model evaluator is the same model family used in the feedback loop, this result cannot be read as evidence of human-perceived quality.","section":"§4.4, Figure 5"},{"comment":"The paper's own Discussion concedes that 'model-generated feedback often diverges from human evaluations.' This is a direct acknowledgment that the load-bearing evaluation proxy is not human-aligned, yet the main claims in the Abstract and Section 4.2 are framed without this caveat. The manuscript should either validate the LLaVA-NeXT-Video evaluator against human judgments (e.g., correlation, calibration, or agreement on held-out videos) or explicitly restrict all claims of improvement to model-assessed metrics and treat the human results as the primary evidence, which currently does not support iterative improvement.","section":"§5 (Discussion and Limitation)"}],"minor_comments":[{"comment":"Four of the nine seed papers (KNOWNET, TUNING, THREADS, DYNAMIC) are authored by members of the authors' research group. This is a potential source of bias and should be disclosed and discussed in the paper.","section":"§4 (Seed Paper Selection)"},{"comment":"The figure caption states that all score axes are standardized to a range between 1.75 and 4.75, which is unusual for a 1–5 scale and may exaggerate small differences; the paper should report the original score scales and the standardization method.","section":"Figure 4 caption"},{"comment":"The terms 'Flashtalk' and 'flash-talk' are used inconsistently, and the 'Flashtalk Feedback Agent' metrics listed in the Appendix (Clarity, Curiosity, Effectiveness) do not match the single metric per agent mentioned in Section 4; this inconsistency should be resolved.","section":"§3.1 and §3.2"},{"comment":"The 'Type Evaluation' column in Table 2 is ambiguous and appears to be a typo; the table would be clearer with an explicit column for agent type.","section":"Appendix, Table 2"},{"comment":"The reference for MoviePy (Zulko, 2014) cites a GitHub URL 'charlespwd/project-title' rather than the MoviePy repository; this should be corrected.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a preprint under review, and the title 'Stealing Creator's Workflow' is informal for a journal submission but not a blocking issue. The most serious concern is the circularity of the evaluation: the same MLLM family provides feedback, guides prompt revision, and then rates the final videos. The authors' own human data contradict the model-based improvement claim, so the revision will need substantial new evaluation work, not just presentational changes. I also note the undisclosed use of four papers authored by the authors' group and the lack of inter-annotator reliability in the human study; these should be addressed in the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"SciTalk is worth knowing about: it's the first system I know of that generates short-form scientific videos from paper text, figures, and avatars using a multi-agent pipeline with an iterative MLLM feedback loop. The system is described in enough detail to be reimplemented, and the authors are honest about its limitations. But the headline claim—that the iterative loop improves accuracy and engagement—is not actually supported by the evidence in the paper.\n\nThe problem is structural. The Feedback Agents and the Evaluation Agent are all LLaVA-NeXT-Video. The Reflection Agent rewrites prompts to raise the scores from that model family, and then the same model family evaluates the final videos. So the upward trends in Figures 3a and 4 partly just show the loop converging to the evaluator's preferences. The paper's own human data in Section 4.3 go the other direction: Scene Readability drops from 3.8 to 2.2 and Audio-Visual Alignment from 4.2 to 2.7 across iterations, while model scores rise. The authors concede in Section 5 that model feedback 'often diverges from human evaluations.' That concession is accurate, and it undermines the abstract's claim of outperforming simple prompting, since the comparison against the single-agent baseline (Figure 4) is also model-only.\n\nOther soft spots are more conventional: the multi-agent vs. single-agent comparison is presented only as figure trends with no numeric values or significance tests; human evaluation uses two internal annotators without inter-annotator agreement; four of the nine seed papers come from the authors' own group; and 'invalid outputs were excluded' without stated criteria. None of these are fatal for a first system paper, but they need fixing.\n\nWhat I like: the agent decomposition is sensible—separating script, scene plan, text, layout, and effects—and the sub-scene feedback design is a practical response to MLLM token limits. The sanity-check mechanism, even though disabled in the main experiments, shows the authors know where their agents fail. The qualitative examples are informative. This is a genuine integration effort, not a toy.\n\nWho is this for? Anyone building automatic science communication or paper-to-video tools. It deserves peer review, but it should come back with a real human evaluation: more raters, agreement metrics, and ideally a held-out validation of the MLLM evaluator against human judgments—or evaluation by a different model than the one giving feedback. Numeric results for the baseline comparison and release of code/data/videos would also help.\n\nVerdict: conditional accept, not reject.","headline":"SciTalk is a genuinely new pipeline for paper-to-video generation, but its headline improvement claim is scored by the same MLLM that drives the loop, and the paper's own human data undercut it.","tokens_in":17851,"tokens_out":2210,"would_cite":false,"duration_ms":22110,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SciTalk, a multi-agent pipeline that composes short-form videos from a paper's own text and figures, claims to outperform single-prompt generation for scientific accuracy and engagement.","keywords":["agentic framework","scientific short-form video","multi-LLM agents","iterative feedback","vision-language feedback","flash talk","grounded video generation","prompt refinement"],"falsifier":"Take a fresh set of papers, run the five-iteration SciTalk loop on each, and have independent human raters score every iteration with the paper's own rubric; if human scores for scene readability and audio-visual alignment decline while the model's scores rise, as the paper's own numbers already suggest (SR 3.8 to 2.2; AVA 4.2 to 2.7), the claim that iterative feedback improves video quality is falsified.","tokens_in":16830,"feed_emoji":"🎬","tokens_out":6020,"duration_ms":55369,"temperature":0.7,"pith_summary":"The paper proposes SciTalk, a fully automatic pipeline that turns a research paper into a short-form video by assigning the work to specialized LLM agents—one writes the flash-talk script, one plans sub-scenes, others choose backgrounds, on-screen text, effects, and layouts—and then composing the video from the paper's own text, figures, and screenshots with a video editing library instead of a generative video model. The central claim is that this multi-agent structure produces videos that are more scientifically accurate and engaging than a single-prompt, one-pass baseline. The paper further claims that an iterative feedback loop, in which vision-language agents rate sub-scenes and reflection agents rewrite the generation prompts, progressively improves the system's own quality scores. A sympathetic reader would care because it stakes out a viable alternative to end-to-end generative video for science communication: grounded editing rather than synthetic generation. The authors are careful that results remain below human creator quality and that model-driven 'improvement' can diverge sharply from what human viewers perceive.","feed_headline":"Multi-agent loop turns papers into videos that beat plain prompting","feed_subtitle":"Five iterations of feedback lift model-rated engagement and accuracy, though humans still prefer creator-made videos.","key_machinery":"The load-bearing mechanism is the prompt-refinement feedback loop. Three feedback agents—Flashtalk, Sceneplan, and Text—each powered by the vision-language model LLaVA-NeXT-Video, rate sub-scenes on role-specific metrics such as curiosity, visual relevance and clarity, and key information coverage. Reflection Agents then rewrite the corresponding generation agent's prompt, keeping only feedback relevant to that agent's domain, and the loop repeats for five iterations. Because the model cannot ingest full videos, evaluation happens sub-scene by sub-scene, and a separate Evaluation Agent scores the final assembled video. The loop is what the paper credits for progressive quality gains, and it is also the component whose validity is most in question.","core_discovery":"On its own terms, the paper's discovery is that creator-style workflow decomposition transfers to automatic scientific video production: a coordinated set of agents that plan, edit, and critique can outperform a single prompted generation pass on compositional quality, coherence, and alignment with the scientific message. SciTalk keeps every visual grounded in the source paper, uses a non-generative composition step (MoviePy) to avoid diffusion artifacts, and closes the loop with LLaVA-NeXT-Video-based feedback agents that score sub-scenes and reflection agents that fold the feedback into revised prompts. Over five iterations, model-rated engagement rises and peaks at the fourth iteration, and content accuracy peaks at the third, while human raters show only minor fluctuations with significant declines in scene readability (3.8 to 2.2) and audio-visual alignment (4.2 to 2.7). In a direct comparison on three papers, human evaluators preferred creator-made videos, while the model evaluator favored SciTalk on some clarity and sync metrics—evidence that the claimed gains hold for the automated evaluator but not yet for human perception.","pith_inferences":["The divergence between model and human scores suggests the loop is optimizing for what a vision-language model finds engaging, not what viewers find engaging; a cheap test would be to have the same evaluator judge a set of videos whose viewer retention is already known and see whether its scores predict retention at all.","The recurring visual clutter and overlapping text, despite explicit feedback, suggests prompt-only refinement is a weak lever over layout; a stronger extension would route rendered frames back to the Layout Allocator as visual input for closed-loop correction.","The same four-stage scaffold could generalize to other grounded communication products—conference slides, infographics, or narrated figure walkthroughs—where the bottleneck is also keeping generated content faithful to a source document."],"forward_implications":["If correct, an automatic SciTalk pipeline can produce paper-grounded short-form videos without generative video models, avoiding diffusion artifacts while keeping figures and text sourced from the paper.","The iterative loop can be run any number of times, and the paper's data indicate that model-rated engagement peaks at iteration 4 and content accuracy at iteration 3, so there is an optimal stopping point rather than monotone improvement.","Model-based evaluation can rank SciTalk videos above human creator videos on clarity and sync metrics, meaning automated evaluation alone is not enough to certify quality for a human audience.","Human praise of creator videos over SciTalk videos in direct comparison implies the framework's current ceiling is below skilled human production, so the practical near-term use is assisted drafting rather than replacement."],"supporting_citations":[{"why":"Supplies LLaVA-NeXT-Video, the vision-language model that powers all feedback agents and the Evaluation Agent.","marker":"Zhang et al. (2024)"},{"why":"Supplies GPT-4o, the backbone model for every generation agent whose prompts the feedback loop rewrites.","marker":"OpenAI et al. (2024)"},{"why":"Supplies MoviePy, the video editing library that composes the final clips from paper assets instead of generative video.","marker":"Zulko (2014)"},{"why":"Empirical account of creators' iterative performance and production practices that the four-stage workflow is modeled on.","marker":"Klug (2020)"},{"why":"Documents creators' planning-production-editing behaviors on algorithmic platforms, grounding the creator-workflow inspiration.","marker":"Choi et al. (2023)"},{"why":"Describes creators' algorithmic dependencies and iterative adjustment, motivating the feedback-loop design.","marker":"Herman (2023)"},{"why":"Shows short-form platforms function as informal learning spaces, motivating the scientific-dissemination goal.","marker":"Ghosh & Figueroa (2023)"}],"fun_headline_variants":["Creator-style feedback loop lifts video metrics, not human taste","Automated feedback loop boosts AI-rated video quality, humans disagree","Iterative agent loop: model likes it, humans don't yet","Creator-inspired AI video loop beats plain prompts—on AI metrics","Feedback-driven video agents: AI says better, humans say no"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the vision-language model's ratings of video quality track what human viewers would say, so that prompt changes that raise model scores genuinely improve the videos; the paper's own human data show large divergences, with human scores for scene readability and audio-visual alignment falling across iterations while model scores rose.","fun_headline_variants_meta":{"raw":{"variants":["Creator-style feedback loop lifts video metrics, not human taste","Automated feedback loop boosts AI-rated video quality, humans disagree","Iterative agent loop: model likes it, humans don't yet","Creator-inspired AI video loop beats plain prompts—on AI metrics","Feedback-driven video agents: AI says better, humans say no"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000856,"raw_usage":{"total_tokens":3723,"prompt_tokens":952,"completion_tokens":2771,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":568,"completion_tokens_details":{"reasoning_tokens":2685}},"tokens_in":568,"tokens_out":2771,"duration_ms":18911,"temperature":1.0,"reasoning_tokens":2685,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:08:38.224821+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a fresh set of papers, run the five-iteration SciTalk loop on each, and have independent human raters score every iteration with the paper's own rubric; if human scores for scene readability and audio-visual alignment decline while the model's scores rise, as the paper's own numbers already suggest (SR 3.8 to 2.2; AVA 4.2 to 2.7), the claim that iterative feedback improves video quality is falsified.","supporting_citations":[],"review_version":1}