{"id":"db61441e-8ac7-472c-8c58-5101f0ddb9c8","arxiv_id":"2508.09632","paper_version":6,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Preacher is an agentic pipeline that decomposes a research paper into planned key scenes and generates a coherent video abstract segment by segment.","lead":"This paper introduces Preacher, an AI system that turns a research paper into a video abstract by first summarizing the paper into a scene-by-scene plan and then generating video segments from that plan. It targets faster, more accessible science communication, claiming better results than video-generation models alone across five research fields.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central comparative claim is not supported by the accessible text: no metrics, baselines, or fidelity evaluation for P-CoT are visible.","rationale":"The reader's weakest_assumption correctly identifies that P-CoT must preserve paper content and that the video generator must render scenes faithfully. I agree that this is the load-bearing condition for the central claim. My review adds that the accessible evidence cannot even begin to test this condition: the abstract gives no metrics, baselines, or evaluation protocol, and the supplied full text is not machine-readable due to encoding corruption. This is not a proven flaw in the system, but it means the correctness risk remains unknown. Because no internal contradiction or decisive counter-evidence is available, the appropriate verdict is unchanged: UNVERDICTED, with confidence low. If the actual evaluation contains rigorous factual-fidelity comparisons, the concern would be resolved; if not, the comparative claim should be substantially weakened.","tokens_in":20508,"tokens_out":1915,"duration_ms":25256,"concrete_test":"Obtain the actual PDF and inspect the evaluation section. Specifically: (1) Confirm whether Preacher is compared against end-to-end video generation models on the same papers using any quantitative metric or human study, rather than only qualitative examples. (2) Run a factual-fidelity audit on a sample of generated video abstracts: extract ground-truth key claims from N=20 source papers, then measure recall and precision of those claims appearing in the generated videos. If P-CoT's scene plans preserve content, claim recall should be high (e.g., >0.85) and claim precision should be high (e.g., <10% hallucinated claims). If no such evaluation exists, the 'expertise beyond current video generation models' claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that Preacher is the first paper-to-video agentic system and that it 'successfully generates high-quality video abstracts across five research fields, demonstrating expertise beyond current video generation models.' For this to hold, the top-down planning stage (P-CoT) must preserve the paper's substantive content, and the bottom-up video generator must faithfully render the planned scenes. The abstract describes the architecture but provides no quantitative evidence, no baseline comparison, no human evaluation, and no code. The supplied full text is unreadable due to encoding corruption, so the evaluation section cannot be audited from the available material. What remains is a self-reported qualitative claim. The load-bearing risk is semantic drift: P-CoT iteratively refines key scenes, and if it omits, distorts, or hallucinates claims while preserving fluent visual coherence, the video abstract will look correct but be scientifically wrong. Since the paper's contribution over end-to-end video generation models depends on this cross-modal fidelity, the abstract-level evidence is insufficient to establish the claimed superiority. The concern is not an internal inconsistency; it is that the central empirical assertion is unverified in the accessible record.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Preacher addresses paper-to-video generation by replacing monolithic end-to-end generation with a two-stage agentic pipeline: a top-down phase that decomposes, summarizes, and reformulates the paper into key scenes, supported by a Progressive Chain of Thought (P-CoT) planner, followed by bottom-up generation of video segments that are synthesized into a coherent abstract. The abstract claims this is the first paper-to-video agentic system and that it produces high-quality video abstracts across five research fields, 'demonstrating expertise beyond current video generation models.' Code is promised at a public repository. The submitted full text, however, is almost entirely unreadable due to character-encoding corruption; only the abstract and a few table-like glyphs are recoverable, so the technical description, evaluation protocol, and numerical results cannot be checked from the material provided.","tokens_in":20736,"tokens_out":4449,"duration_ms":54818,"significance":"If the claimed results hold, the paper would make a useful contribution to document-to-video and scientific communication, and the key-scene/P-CoT decomposition is an intuitively plausible way to address the limited context and rigid duration constraints of end-to-end video models. The planned code release is a positive feature. However, the central comparative claim—'expertise beyond current video generation models'—is empirical and requires metrics, baselines, and a fidelity evaluation. None of that evidence is visible in the readable portion of the manuscript, and the body cannot be audited in its current form. The contribution is therefore potentially significant but currently unverified.","major_comments":[{"comment":"The abstract asserts that Preacher 'successfully generates high-quality video abstracts across five research fields, demonstrating expertise beyond current video generation models.' No quantitative results, baseline comparisons, user study, or error analysis are reported in the abstract, and the supplied full text is corrupted beyond readability. This comparative statement is load-bearing: it is the basis for claiming superiority over existing models. Please supply a readable manuscript with explicit metrics, named baselines, evaluation protocols, and statistical support.","section":"Abstract"},{"comment":"The design assumes that the top-down plan preserves the paper's substantive content and that the bottom-up generator faithfully renders the planned scenes. The paper does not provide evidence against semantic drift, omission, or hallucination in P-CoT, nor does it define a fidelity metric connecting generated scenes to the source paper. Since the claimed advantage over end-to-end generation depends on this cross-modal fidelity, a concrete evaluation is necessary—for example, claim-level precision/recall against the source paper, expert annotation of scientific correctness, or human comparison with end-to-end baselines. This is not a circularity objection; it is a request for external grounding.","section":"P-CoT / key-scene fidelity"},{"comment":"The body consists almost entirely of replacement characters; sections, equations, figures, and tables cannot be audited. Table-like structures appear toward the end, but no readable caption, legend, or numeric entries can be extracted. Consequently, the claimed results across 'five research fields' and any comparison with current video-generation models are not independently verifiable from the submission. The manuscript must be regenerated with a valid encoding before a technical review can proceed.","section":"Full text / evaluation tables (passim)"}],"minor_comments":[{"comment":"'Topdown' should be 'top-down.' Also, 'Progressive Chain of Thought (P-CoT)' should be expanded at first use in the readable text; the abstract uses the abbreviation without definition.","section":"Abstract"},{"comment":"All figures and tables need readable captions, axis labels, legends, and numeric values; the current submission does not allow extraction of any quantitative result.","section":"Figures and tables"},{"comment":"The paper should state exactly which 'current video generation models' are compared, which versions are used, and how output quality is measured (automated metrics, human ratings, or both).","section":"Comparison protocol"},{"comment":"If code is to be released, consider including a versioned repository link, configuration details, prompt templates, and example outputs in an appendix so the agentic pipeline can be reproduced.","section":"Reproducibility"}],"recommendation":"uncertain","confidential_remarks":"The submitted PDF is essentially unreadable due to encoding corruption; a normal technical review cannot be performed on the available material. I recommend asking the authors for a corrected PDF or LaTeX source, together with the evaluation section and any supplementary materials, before assigning a substantive verdict. The novelty claim ('first paper-to-video agentic system') should also be checked against prior document-to-video and video-abstract literature once the text is readable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I read the abstract and the corrupted full-text dump. Punchline: this is a plausible system paper with a sensible architecture, but the headline claim of superiority over current video generation models is not backed by anything I can inspect. The full text is encoding garbage, so the evaluation section, figures, and references are unavailable. Treat the empirical claims as unverified.\n\nWhat's genuinely new: the orchestration pattern. Using an LLM to decompose a paper into key scenes, refining via Progressive Chain of Thought, then generating bottom-up segments is a reasonable way to handle long documents and fixed context windows. I can believe that produces more coherent video abstracts than a single text-to-video pass. The five-field demonstration, if real, is a useful proof of concept. The authors deserve credit for identifying the right failure mode—cross-modal semantic drift—and putting P-CoT in the loop to address it.\n\nSoft spots, in proportion: first, no metrics anywhere in the abstract. No baselines, no human evaluation, no error analysis. 'Demonstrating expertise beyond current video generation models' is a comparative claim with no comparator. The 'first' claim also needs a literature check; the abstract cites nothing. Second, the fidelity problem is the load-bearing risk. If P-CoT drifts from the paper's actual claims, the video will look fluent and be wrong. Nothing in the accessible text checks faithfulness to the source. Third, the evaluation might be self-referential—key scenes and P-CoT are author-defined, and if the quality ratings are author-run, that's circular. I'm not saying it is; I'm saying the abstract doesn't rule it out.\n\nWho this is for: researchers in automated science communication and video-generation agents. It's an applied tool, not a scientific contribution, and its value depends entirely on whether the full paper has a real evaluation. If it does, this is a decent system paper. If not, it's a demo.\n\nRecommendation: send to peer review, but with a mandate that reviewers verify the evaluation is real—metrics, baselines, human raters, and code. Desk rejection would be premature; accepting the comparison at face value would be wrong.","headline":"Sensible agentic paper-to-video pipeline, but the comparative claim is unverified and the supplied full text is corrupted.","tokens_in":21235,"tokens_out":2681,"would_cite":false,"duration_ms":29374,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Preacher converts research papers into structured video abstracts by planning first and rendering second.","keywords":["paper-to-video","video abstract","agentic system","progressive chain of thought","key scene planning","cross-modal generation","text-to-video","scientific communication"],"falsifier":"Run Preacher on a paper containing precise quantitative claims outside the five tested fields, then have domain experts mark whether every central claim, method step, and conclusion appears in the final video. If the video is visually fluent but misses or distorts a stated result, the pipeline's fidelity claim fails.","tokens_in":20394,"feed_emoji":"🎬","tokens_out":2368,"duration_ms":29681,"temperature":0.7,"pith_summary":"This paper claims that the task of turning a research paper into a video abstract is better solved by an agentic pipeline than by prompting a video-generation model directly. It proposes Preacher, which decomposes the paper into a structured plan, refines that plan into key scenes through a Progressive Chain of Thought, and then renders and stitches scenes into a coherent video. The claim is that this top-down planning plus bottom-up generation yields faithful, high-quality video abstracts across five research fields, going beyond what current video models can do alone. A sympathetic reader would care because it reframes the bottleneck from rendering to planning: the hard part of paper-to-video is deciding what to show, not generating pixels.","feed_headline":"Agentic pipeline turns research papers into video abstracts","feed_subtitle":"Planning scenes before rendering lets domain-specific content survive where direct video generation falls short.","key_machinery":"The load-bearing mechanism is the division between top-down semantic planning and bottom-up visual synthesis, joined by 'key scenes' as the cross-modal unit of alignment. P-CoT (Progressive Chain of Thought) is the iterative planner that refines these scenes granularly, so that each planned scene encodes the paper's concepts, methods, and conclusions before any video generation happens. The planner, not the generator, is what carries the paper's expertise.","core_discovery":"The paper's central claim is that paper-to-video should be treated as an agentic, multi-stage process rather than a single end-to-end generation call. Preacher first reads the paper and decomposes, summarizes, and reformulates its content into a set of key scenes; a Progressive Chain of Thought (P-CoT) then refines these scenes iteratively to align the textual and visual representations. After the scene plan is settled, videos are generated segment by segment from the bottom up and synthesized into one coherent abstract. The authors assert that this design succeeds in producing high-quality video abstracts across five research fields and demonstrates expertise beyond current video generation","pith_inferences":["Editorial inference: the same planning-then-rendering split could extend beyond papers to other long structured documents—textbooks, patents, clinical guidelines, or technical reports—where the limiting factor is faithful content selection rather than visual fluency.","Editorial inference: the most likely failure mode is silent drift in the P-CoT scene plan, where the generated video is fluent and visually coherent but omits or distorts a central result; a verification layer that checks scenes against the source paper would be the natural next step.","Editorial inference: key scenes could be reused as a lightweight evaluation artifact—a human or automated checker could score faithfulness of the plan alone, before spending compute on video rendering.","Editorial inference: because the output is assembled from segments, the system may be more controllable than end-to-end generation, but it also inherits the risk of inconsistent transitions, so scene stitching is a hidden quality bottleneck the paper does not foreground."],"forward_implications":["If the claim holds, research papers can be turned into watchable video abstracts without retraining or fine-tuning a video generation model on scientific content.","The approach separates content planning from rendering, so improvements in video generators can be absorbed by swapping the renderer while keeping the planning layer intact.","Because the scene plan is explicit, the system can in principle expose which parts of a paper were selected for the abstract, making the summarization process inspectable rather than a black box.","The demonstrated breadth across five research fields suggests the pipeline is not tied to one visual style or domain vocabulary, assuming the planning layer generalizes.","The central comparison to direct video generation implies that the advantage comes from the planning stage, which is a testable and reusable component on its own."],"supporting_citations":[],"fun_headline_variants":["Preacher uses agentic planning to turn papers into videos","Paper-to-video agent breaks generation into planned scenes","Iterative scene planning lets AI turn papers into videos","From paper to video: agentic decomposition beats direct gen","Preacher's P-CoT aligns text and video for better abstracts"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that the planned key scenes faithfully preserve the paper's substantive content, so errors or omissions introduced during planning cannot be corrected later by the video generator.","fun_headline_variants_meta":{"raw":{"variants":["Preacher uses agentic planning to turn papers into videos","Paper-to-video agent breaks generation into planned scenes","Iterative scene planning lets AI turn papers into videos","From paper to video: agentic decomposition beats direct gen","Preacher's P-CoT aligns text and video for better abstracts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000212,"raw_usage":{"total_tokens":1215,"prompt_tokens":668,"completion_tokens":547,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":412,"completion_tokens_details":{"reasoning_tokens":466}},"tokens_in":412,"tokens_out":547,"duration_ms":6444,"temperature":1.0,"reasoning_tokens":466,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:55:03.152356+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run Preacher on a paper containing precise quantitative claims outside the five tested fields, then have domain experts mark whether every central claim, method step, and conclusion appears in the final video. If the video is visually fluent but misses or distorts a stated result, the pipeline's fidelity claim fails.","supporting_citations":[],"review_version":1}