{"id":"1f17c5b7-ce5e-44a0-8f84-913e57c6c2d9","arxiv_id":"2607.04438","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A five-skill agent pipeline generates an editable poster, video deck, and bilingual blog from a paper PDF, binds them in an interactive viewer, and reports poster scores above the authors' own under two VLM judges.","lead":"ResearchStudio-Reel turns one accepted paper PDF into three editable research-dissemination artifacts — a PowerPoint poster, a narrated talk video, and a bilingual Word blog — and links them in one interactive viewer. On a 100-paper benchmark, its posters score higher than the authors' own posters on aesthetics under two AI judges, though the top-line numbers rest on subjective machine ratings.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Aesthetic superiority over author posters rests on VLM judges from the same model families as the generators and on a pipeline explicitly optimized to those judges; no human or third-family validation exists.","rationale":"The reader's weakest assumption identifies the same load-bearing issue: the two VLM judges are from the same families as the generators and may systematically prefer their own generative style. My analysis agrees and sharpens it: the system is not merely scored by these judges, it is explicitly optimized to them via the measured-fill loop and composition axes. This makes the 3.56 vs. 3.03 aesthetic win and the 74/95 per-paper win rates plausibly a Goodhart artifact rather than a genuine quality advantage. The paper is transparent about the proxy nature of the aesthetic scores and about the absence of human raters, which is why the verdict should remain CONDITIONAL rather than escalate to REJECT. The proposed concrete test — a third-family VLM judge — directly tests whether same-family bias explains the win rates. If the win rate holds under Gemini-3.1 Pro, the concern is weakened; if it collapses, the abstract's headline comparison is unsupported. The proxy-based model access in Appendix C is an additional reproducibility concern but is secondary to the judge-validity issue. Overall, the paper's engineering contributions and honest disclosures are real, but the central quantitative claim should be read as conditional on independent validation.","tokens_in":26096,"tokens_out":7208,"duration_ms":81070,"concrete_test":"Score all 100 Paper2Poster posters (Reel Claude Code, Reel Codex, author ground-truth, and all baselines) with a third VLM judge from a disjoint family — e.g., Gemini-3.1 Pro — using the identical six-criterion rubric, downscaling, and true-content-size rendering protocol. Compute the per-paper overall win rate of Reel (Claude Code) against author ground-truth. If the win rate drops from the reported 85%/100% (of non-tied papers) to near chance (e.g., ≤55%), the original win rates are inflated by judge–generator family correlation and the abstract's 'exceeds the authors' posters' claim does not generalize beyond the original two judges.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim — that ResearchStudio-Reel (Claude Code) exceeds the authors' posters in average aesthetics (3.56 vs. 3.03) and wins overall on 74/100 and 95/100 papers (§3, Table 1) — depends on the validity and impartiality of the two VLM judges. That condition is least secure for two reasons. First, the Claude Code configuration is generated by claude-opus-4.8 and scored by claude-opus-4.8, and the Codex configuration is generated by gpt-5.5 and scored by gpt-5.5; a VLM can systematically prefer layouts that resemble its own generative style, so the head-to-head against human posters is not independent. Second, the pipeline's measured-fill loop and composition axes are explicitly engineered to maximize the exact six-criterion rubric used by these judges (A2–A5, §2.2), so the system is optimized to the same metric on which it is evaluated. The paper itself concedes this: 'aesthetic appeal is inherently subjective... the aesthetic numbers should be read as a proxy signal rather than proof' (§3 Analysis), and a third-party human-rater study is 'left for future work' (Appendix A). Additionally, Appendix C discloses that all model runs went through a private Copilot API proxy, which makes independent reproduction of the exact judge harder. Thus the 'exceeds the authors' posters' claim is, as of now, an in-family, metric-optimized result rather than evidence about human aesthetic preference. This is a correctness-risk issue for the paper's most load-bearing assertion, not a minor caveat.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes ResearchStudio-Reel, a skill-based pipeline that takes a paper PDF through a shared extraction stage (Paper2Assets) and produces a PowerPoint poster, a narrated video with editable deck, a bilingual Word blog, and a Paper2Reel interactive viewer that binds poster, video, and blog at the section level. The main empirical claim is on the Paper2Poster benchmark: under two VLM judges, the Claude Code configuration achieves 3.56 average aesthetics versus 3.03 for author posters, wins overall on 74/100 and 95/100 papers, and is best among automated systems on the aesthetic criteria. The paper also contributes deterministic release gates, a measured-fill poster loop, operational profiling, and capability audits for video and blog. Video and blog quality are not quantitatively evaluated, and the paper explicitly frames the aesthetic numbers as a VLM proxy signal.","tokens_in":26417,"tokens_out":7398,"duration_ms":84607,"significance":"If the poster-quality result were independently confirmed, the system would be a substantial engineering advance: one shared asset bundle feeds native-editable PowerPoint and Word artifacts, aligned through an interactive convergence layer, with deterministic package gates that make the delivery contract checkable. Strengths of the evaluation include the external Paper2Poster benchmark, true-size rendering before scoring, two judges from different model families, per-paper win-rate computation, a clear discussion of the PaperQuiz/aesthetics tension, and unusually candid limitations. The central weakness is that the headline comparison is in-family: the primary configuration is generated by claude-opus-4.8 and scored by claude-opus-4.8, the Codex configuration by gpt-5.5 and scored by gpt-5.5, and the pipeline is explicitly optimized to the rubric used by these judges. The paper's own caveats are appropriate but do not overcome the evidentiary gap.","major_comments":[{"comment":"The headline claim that ResearchStudio-Reel (Claude Code) 'exceeds the authors' posters in average aesthetics (3.56 vs. 3.03)' and wins 74/100 and 95/100 is evaluated exclusively by two VLM judges, claude-opus-4.8 and gpt-5.5. The primary configuration is generated by claude-opus-4.8 and scored by claude-opus-4.8; the Codex configuration is generated by gpt-5.5 and scored by gpt-5.5. Moreover, the measured-fill loop and the composition axes in §2.2.2 (A2–A5) are explicitly engineered to satisfy the same six-criterion rubric used by these judges. The paper's disclaimer in §3 Analysis that the aesthetic numbers are 'a proxy signal rather than proof' is appropriate, but the abstract and conclusion state the result without that nuance. This is load-bearing: the claim is currently an in-family, metric-optimized result. Please add a third-party human-rater study, or at minimum a third-family j","section":"§3, Table 1; §2.2"},{"comment":"The comparison against prior automated systems and the author ground-truth depends on baselines that are 'reproduced with best efforts and scored by us under Claude Code with claude-opus-4.8.' Because the prior systems are stochastic agentic pipelines, 'best efforts' is not a controlled condition; as reported, the 'best among automated systems' and the win-rate claims cannot be independently checked. Please release per-paper scores and generated poster files for all systems, or use the original benchmark's reported scores where available, and specify the exact version and configuration used for each reproduced baseline.","section":"§3, Table 1 note †"},{"comment":"The paper's title and contribution list promise poster, video, and blog, but quantitative evaluation is limited to posters. Tables 2 and 3 are feature checklists, and Appendix A explicitly states that video and blog are not evaluated and that no human editing-effort or navigation-understanding study exists. The 'last mile automation' claim for the full workspace is therefore supported only by existence checks. Please add targeted quantitative evaluations for video (duration accuracy, caption/timeline alignment, visual-cue correctness) and blog (bilingual fact consistency, figure placement, layout defects), or substantially soften the abstract and conclusion claims to match the evidence.","section":"§3, 'Capability coverage'; Appendix A"},{"comment":"All reported runs went through a private Copilot API proxy, and the model identifiers 'name the underlying models as served by that proxy.' This makes exact reproduction of the benchmark numbers impossible and leaves open the possibility that proxy-specific behavior contributed to the scores. Please provide the exact proxy configuration/version and, for the headline numbers, run the same pipeline on first-party endpoints or release the full run logs.","section":"Appendix C"}],"minor_comments":[{"comment":"The caption says 'claude-opus-4-8' while the rest of the paper uses 'claude-opus-4.8'; please harmonize.","section":"Table 4 caption"},{"comment":"The system's skill is named Paper2Video, which is also the name of the external benchmark/system in reference [4]. This naming collision is confusing in the related-work discussion; please disambiguate, e.g., by referring to the external work as 'Paper2Video [4] (benchmark)'.","section":"§5.2"},{"comment":"The figure includes 'Claude Code (claude-opus-4.7)' and 'Claude Code (claude-opus-4.6)' panels, but these model versions are not listed in Table 1 and are not described in the text. Please state what these panels illustrate and why they are not part of the quantitative comparison.","section":"Figure 10"},{"comment":"The win rates are reported as 74/100 and 95/100 with percentages of non-tied papers. Please also report the number of ties and losses so the head-to-head distribution is fully specified.","section":"§3, per-paper win rate"},{"comment":"The Reader-Reconstruction Preference gate is described as an optional in-loop signal but no result using it is reported. Please state whether it was enabled for the Table 1 poster runs, and if so, what effect it had.","section":"§2.2.4, RRP gate"}],"recommendation":"major_revision","confidential_remarks":"The engineering contribution is substantial and the paper is unusually candid about its limitations. My main reservation is that the only quantitative headline result is an in-family VLM evaluation of aesthetic quality, so the 'exceeds the authors' posters' claim is not yet established as evidence about human preference. I would be willing to accept a revision that adds a human-rater or third-family judge study for posters, releases the baseline reproduction details, and provides at least minimal quantitative validation of the video and blog artifacts."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a genuine engineering contribution — a five-skill pipeline (shared extractor, poster, video, blog, and a convergence viewer) that turns one PDF into editable PPTX poster, PPTX video deck, bilingual DOCX blog, and a navigable HTML surface. The release gates are unusually concrete: categorical fill verdicts, circuit breaker, figure-fill checks, deterministic package validation. That part is well done and worth reading closely if you build agentic document pipelines.\n\nThe soft spot is the quantitative headline. The paper claims its posters exceed the authors' posters on aesthetics (3.56 vs 3.03) and win on 74/100 and 95/100 papers. That's under two VLM judges, claude-opus-4.8 and gpt-5.5. The Claude Code configuration is generated by claude-opus-4.8 and the Codex configuration by gpt-5.5, so the judge is from the same model family as the generator for each configuration. Worse, the measured-fill loop is explicitly engineered to hit the same six-criterion rubric those judges use. So the result is an in-family, metric-optimized comparison, not evidence about human preference. The paper fesses up to this — it says the numbers are a proxy signal, not proof, and leaves a human-rater study to future work. That honesty counts, but it means the abstract's 'exceeds the authors' posters' should carry a heavy caveat.\n\nOther soft spots: PaperQuiz ordering inverts — Reel scores below P2P and single-shot GPT-5.5 on comprehension, which the paper discusses. Video and blog have no quantitative evaluation, only capability audits. Baselines were reproduced by the authors' own team, and all model runs went through a private Copilot API proxy, which makes independent replication of the exact judge harder. Also worth noting: concurrent OmniPresent shares co-authors but isn't discussed as such.\n\nProportionate verdict: the engineering is solid and the evaluation is honest but under-powered for the headline. I'd send this to peer review, with the expectation that the reviewers ask for a third-party human rater study, a judge from a different family than any generator, and a softer abstract. If those land, this could be a useful reference for editable multi-artifact dissemination.","headline":"A carefully engineered paper-to-poster/video/blog pipeline whose headline superiority over human posters is real only under two VLM judges from the same model families; valuable as an engineering contribution, not as evidence of human aesthetic preference.","tokens_in":27059,"tokens_out":3529,"would_cite":true,"duration_ms":39677,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims a single pipeline turns one paper PDF into an editable poster, narrated video, and bilingual blog in one navigable viewer, with posters scoring above the authors' own on aesthetics under two automated judges on 100 papers.","keywords":["paper-to-poster","paper-to-video","paper-to-blog","dissemination automation","measured-fill loop","section-level alignment","native-editable deliverables","vision-language model evaluation"],"falsifier":"Have an independent panel of human raters, blind to provenance, score the same 100 poster pairs with the same six-criterion rubric; if the authors' own posters win on aesthetics or overall quality, the central claim collapses. A cheaper mechanistic check: cross-score outputs, letting the judge used in the system's own configuration score the other family's outputs and vice versa, and test whether same-family scoring systematically inflates scores.","tokens_in":25899,"feed_emoji":"🖼️","tokens_out":10028,"duration_ms":98560,"temperature":0.7,"pith_summary":"The paper tries to establish that the last mile of research dissemination — turning an accepted paper into a conference poster, a talk video, and a blog post — can be automated as a single editable, navigable workspace, not as three disconnected one-off renders. Its system reads the paper PDF once into a shared asset bundle, then produces an editable PowerPoint poster, a narrated video with an editable slide deck, and a bilingual Word blog, all cross-linked in an interactive viewer that maps poster regions to video segments and blog passages. The quantitative claim is that on a 100-paper benchmark judged by two vision-language models, the automated posters score above the authors' own posters on average aesthetics (3.56 vs 3.03) and win on overall quality on 74 and 95 of the 100 papers under the two judges. A reader should care because the artifacts are native-editable source files: an author can fix a typo, swap a figure, or re-cut a slide without regenerating the whole pipeline.","feed_headline":"Auto posters beat authors' on up to 95 of 100 papers","feed_subtitle":"One PDF becomes an editable poster, narrated video, and bilingual blog, all cross-linked in one viewer.","key_machinery":"Key machinery: the shared asset bundle and the section-level alignment record. The bundle (the paper's text, cleaned figure crops, captions, metadata, a nine-section summary, and narration clips, all with stable section IDs and figure handles) is extracted once and consumed verbatim by all three generators, so the figure in the poster is the figure in the video and blog. The alignment record is a sidecar that maps each canonical section ID to its poster block, slide targets, video timestamps, subtitle tracks, and blog passages; the convergence layer reads this sidecar instead of inferring boundaries from pixels. Supporting these is the poster's measured-fill loop, a categorical controller th","core_discovery":"The paper claims that a five-skill architecture can automate the last mile of dissemination while keeping outputs editable: one shared extractor parses the paper PDF once into a bundle of cleaned figures, captions, metadata, a nine-section summary, and narration clips with stable IDs; three generators consume that bundle verbatim to emit an editable PowerPoint poster, a narrated video with an editable deck and a timeline sidecar, and a bilingual Word blog; and a convergence layer reads the alignment record to bind them into one poster-first viewer. The poster generator's distinctive mechanism is a measured-fill loop that measures each section's fill ratio, maps it to one of five categorical","pith_inferences":["The shared-bundle-plus-alignment pattern is a transferable delivery contract: the same architecture could produce other native-format outcomes (teaching slides, lay summaries, briefing packs) without redesigning the pipeline.","The measured-fill loop generalizes beyond posters: any fixed-canvas layout task (slides, infographics, one-page reports) can be driven by the same five-state categorical controller with one deterministic move per state.","A human reading-and-recall study, which the paper leaves for future work, would settle whether the PaperQuiz-versus-aesthetics inversion reflects a genuine tradeoff or an artifact of exact-match grading; if human recall tracks the aesthetic scores, the fill loop's density target may be miscalibrated."],"forward_implications":["A single extraction pass means the figure in the poster is the same figure in the video and blog; no downstream generator re-crops or re-parses the paper.","The native-editable source files are first-class deliverables, so authors can revise a poster, deck, or blog in PowerPoint or Word without regenerating the other artifacts or losing alignment.","Poster convergence is auditable: the fill loop terminates only when every section sits in the 90–98% band and every figure meets its size floor, and a circuit breaker ships the best-measured state instead of grinding forever.","The video's timeline sidecar keeps sections addressable after export, letting the interactive viewer seek by section rather than infer boundaries from the MP4.","The poster-quality advantage transfers across two different model/harness settings, indicating the workflow rather than a single model carries the result."],"fun_headline_variants":["Auto posters beat authors' on 95 of 100 papers","One PDF becomes editable poster, video, blog, viewer","5 skills automate paper dissemination into editable formats","Paper to poster, video, blog: automated and editable","Auto poster wins on aesthetics, beats authors on 95/100"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The headline comparison assumes the two automated judges' aesthetic ratings are fair and unbiased even though the same model family both generates a configuration and scores the output; the paper itself says these aesthetic numbers are a proxy signal, not proof.","fun_headline_variants_meta":{"raw":{"variants":["Auto posters beat authors' on 95 of 100 papers","One PDF becomes editable poster, video, blog, viewer","5 skills automate paper dissemination into editable formats","Paper to poster, video, blog: automated and editable","Auto poster wins on aesthetics, beats authors on 95/100"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1186,"prompt_tokens":818,"completion_tokens":368,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":562,"completion_tokens_details":{"reasoning_tokens":287}},"tokens_in":562,"tokens_out":368,"duration_ms":4709,"temperature":1.0,"reasoning_tokens":287,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T08:38:09.686297+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have an independent panel of human raters, blind to provenance, score the same 100 poster pairs with the same six-criterion rubric; if the authors' own posters win on aesthetics or overall quality, the central claim collapses. A cheaper mechanistic check: cross-score outputs, letting the judge used in the system's own configuration score the other family's outputs and vice versa, and test whether same-family scoring systematically inflates scores.","supporting_citations":[],"review_version":2}