{"id":"50396fef-bd72-4cce-b45e-7e2899d8a37b","arxiv_id":"2605.15202","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"DeepSlide introduces a multi-agent system for full presentation preparation that matches baselines on slide quality but improves narrative flow, pacing, and script synergy via a new dual-scoreboard benchmark.","lead":"DeepSlide is a human-in-the-loop multi-agent system that plans time-budgeted narratives, retrieves evidence, generates matching slides and scripts, and supports rehearsal with attention cues. A smart generalist might read it to understand how AI can move beyond making pretty slides to improving the actual flow and delivery of talks.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Dual-scoreboard separation of artifact vs. delivery metrics may contain confounding overlap","rationale":"The reader's weakest_assumption matches the identified load-bearing point exactly. The proposed correlation-plus-residual check is a direct, low-cost verification that would either confirm metric independence or show that the headline differential gains are partly artifact-driven. No other internal inconsistency is visible from the given description.","tokens_in":1699,"tokens_out":320,"duration_ms":27838,"concrete_test":"From the evaluation section, extract the exact scoring rubrics and aggregation formulas for both scoreboards; compute Pearson correlation between the two aggregate scores over the 20 domains. If r > 0.25, regress delivery scores on artifact scores and re-compare DeepSlide vs. baselines on the residuals; if the reported larger gains fall below significance, the separation assumption fails.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central empirical claim—that DeepSlide matches baselines on artifact quality but shows larger gains on delivery metrics—depends on the dual-scoreboard benchmark cleanly isolating static slide quality from dynamic aspects (narrative flow, pacing precision, slide-script synergy, attention guidance) without overlap or bias. Delivery metrics such as pacing and narrative flow are plausibly downstream of artifact properties (e.g., clearer slides enable better timing and synergy), so any improvement in slide generation could leak into delivery scores. The abstract provides no description of metric definitions, orthogonality tests, or controls for this dependence, leaving the differential-gain result vulnerable to metric design rather than system capability.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents DeepSlide, a human-in-the-loop multi-agent system for full presentation preparation, integrating a controllable logical-chain planner with per-node time budgets, a lightweight content-tree retriever for grounding, Markov-style sequential rendering with style inheritance, and sandboxed execution for renderability. It introduces a dual-scoreboard benchmark to separate static artifact quality from dynamic delivery excellence and claims that, across 20 domains and diverse audience profiles, DeepSlide matches strong baselines on artifact quality while achieving larger gains on delivery metrics including narrative flow, pacing precision, slide-script synergy, and attention guidance.","tokens_in":1796,"tokens_out":407,"duration_ms":42350,"significance":"If the empirical results hold under a validated benchmark, the work could meaningfully advance AI-assisted scholarly communication by shifting focus from static slide artifacts to the full delivery process. The integration of planning, retrieval, and rehearsal components represents a practical step toward more usable presentation tools, and the dual-scoreboard idea, if shown to be non-confounded, would be a useful methodological contribution for future evaluations in this area.","major_comments":[{"comment":"The central empirical claim—that DeepSlide matches baselines on artifact quality but shows larger gains on delivery metrics—depends on the dual-scoreboard benchmark cleanly isolating static slide quality from dynamic aspects without overlap or bias. The manuscript provides no description of metric definitions, orthogonality tests between the two scoreboards, or controls for potential leakage (e.g., clearer artifact slides enabling better pacing and synergy by construction). This is load-bearing for interpreting the differential-gain result and must be addressed with explicit validation.","section":"Benchmark and Evaluation"}],"minor_comments":[{"comment":"The abstract states performance gains but does not include any quantitative results, baseline details, or statistical tests; moving a concise summary of key numbers into the abstract would improve readability.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback and for recognizing the potential of DeepSlide to advance AI-assisted scholarly communication by emphasizing the full delivery process. We address the major comment on the dual-scoreboard benchmark below.","responses":[{"response":"We agree that explicit validation of the separation is essential to support the differential-gain interpretation. In the revised manuscript we will expand the benchmark description (currently in Section 4) with (i) precise definitions and scoring procedures for every metric on both the artifact and delivery scoreboards, (ii) quantitative orthogonality analysis (Pearson and Spearman correlations across the 20 domains), and (iii) leakage-control experiments that evaluate delivery metrics on fixed baseline artifacts and artifact metrics on fixed scripts. These additions will directly address the concern about confounding and provide the requested explicit validation.","revision_made":"yes","referee_comment":"The central empirical claim—that DeepSlide matches baselines on artifact quality but shows larger gains on delivery metrics—depends on the dual-scoreboard benchmark cleanly isolating static slide quality from dynamic aspects without overlap or bias. The manuscript provides no description of metric definitions, orthogonality tests between the two scoreboards, or controls for potential leakage (e.g., clearer artifact slides enabling better pacing and synergy by construction). This is load-bearing for interpreting the differential-gain result and must be addressed with explicit validation."}],"tokens_in":1317,"tokens_out":297,"duration_ms":33563,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that DeepSlide tries to handle the full presentation process rather than stopping at slide creation, but the abstract gives no numbers to show it works better. It brings together a logical chain planner that assigns time budgets to parts of the narrative, a retriever for grounding content, sequential rendering, and support for rehearsal. The dual-scoreboard benchmark is meant to measure both how good the slides look and how well the delivery goes, like pacing and how the script matches the slides. This setup is new in combining those pieces with explicit attention to delivery metrics and human-in-the-loop elements. It does a solid job describing a practical pipeline that could help people prepare talks more effectively. The weak part is the lack of any actual performance numbers, baseline comparisons, or details on how the benchmark was tested. The claim of larger gains on delivery across 20 domains sounds promising but can't be checked without the data. There's also a risk that the delivery scores are not fully independent from the artifact quality, since better slides might naturally lead to better pacing. This paper would interest people who build or use AI tools for creating presentations in academic or professional settings. A reader working on similar systems could pick up ideas on planning with time constraints or evaluation methods. It has enough of a new angle to deserve a serious referee who can look at the full experiments and methods. I would send it to peer review but ask the authors to add the missing quantitative results and any tests for metric independence right away.","headline":"DeepSlide shifts slide generation toward full delivery support with time-budgeted planning and a dual-scoreboard benchmark, but the abstract supplies no numbers or validation to back the performance claims.","tokens_in":2316,"tokens_out":376,"would_cite":false,"duration_ms":54601,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"DeepSlide's dual-scoreboard and logical-chain planner have no structural overlap with RS","alignment":"orthogonal","rationale":"The paper introduces a multi-agent pipeline for time-budgeted narrative planning, content-tree retrieval, Markov-style rendering, and a dual-scoreboard benchmark separating artifact quality from delivery metrics. None of these components invoke recognition cost J(x), golden-ratio identities, 8-tick periodicity, or any forcing from a single distinction. The work lies entirely in the domain of AI presentation systems and human-in-the-loop evaluation, where RS supplies no predictions or contradictions.","tokens_in":63571,"confidence":"high","tokens_out":138,"duration_ms":8992,"cache_read_input_tokens":32896,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"DeepSlide is a multi-agent system that plans time-budgeted narratives and generates synced slides and scripts to improve delivery while matching visual quality.","keywords":["AI presentation generation","multi-agent systems","narrative planning","slide-script generation","presentation delivery","benchmark evaluation","human-in-the-loop AI"],"falsifier":"A head-to-head user study in which independent raters score the same source content delivered by DeepSlide versus baseline generators and find no advantage or a reversal on narrative flow, pacing precision, or attention guidance scores.","tokens_in":2578,"feed_emoji":"📊","tokens_out":498,"duration_ms":38058,"temperature":0.7,"pith_summary":"Most AI slide tools focus on creating visually plausible decks but leave pacing, narrative structure, and script synergy to the user. DeepSlide addresses this by supporting the full process from requirement gathering through rehearsal, using a logical-chain planner that assigns time budgets to each narrative step. It adds a content retriever for grounding claims, sequential rendering that inherits styles, and sandboxed execution to keep outputs renderable. A new dual-scoreboard benchmark measures static slide quality separately from dynamic delivery aspects such as flow and attention guidance. Across twenty domains the system matches strong baselines on appearance yet shows larger gains on delivery metrics.","feed_headline":"DeepSlide matches slide visuals but lifts narrative flow and pacing","feed_subtitle":"A multi-agent planner assigns time budgets to story nodes and produces matched slides and scripts to improve delivery across twenty domains.","key_machinery":"A controllable logical-chain planner with per-node time budgets that structures the narrative and enforces pacing precision during generation.","core_discovery":"DeepSlide is a human-in-the-loop multi-agent system that supports the full presentation process from requirement elicitation and time-budgeted narrative planning, to evidence-grounded slide-script generation, attention augmentation, and rehearsal support. It integrates a controllable logical-chain planner with per-node time budgets, a lightweight content-tree retriever for grounding, Markov-style sequential rendering with style inheritance, and sandboxed execution with minimal repair to ensure renderability. Evaluation on a dual-scoreboard benchmark across twenty domains shows it matches strong baselines on artifact quality while achieving larger gains on delivery metrics including narrative","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["DeepSlide plans full presentations with time budgeted narratives","DeepSlide lifts delivery while matching visual slide quality","DeepSlide multi-agent planning boosts narrative and pacing precision","DeepSlide generates matched slides and scripts for delivery gains"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The dual-scoreboard benchmark cleanly separates static artifact quality from dynamic delivery excellence without overlap or bias in the evaluation metrics.","fun_headline_variants_meta":{"raw":{"variants":["DeepSlide plans full presentations with time budgeted narratives","DeepSlide lifts delivery while matching visual slide quality","DeepSlide multi-agent planning boosts narrative and pacing precision","DeepSlide generates matched slides and scripts for delivery gains"]},"model":"grok-4.3","cost_usd":0.011047,"raw_usage":{"total_tokens":4778,"prompt_tokens":665,"num_sources_used":0,"completion_tokens":52,"cost_in_usd_ticks":110465500,"prompt_tokens_details":{"text_tokens":665,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":4061,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":665,"tokens_out":52,"duration_ms":65028,"temperature":1.0,"reasoning_tokens":4061,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-19T17:58:19.195575+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A head-to-head user study in which independent raters score the same source content delivered by DeepSlide versus baseline generators and find no advantage or a reversal on narrative flow, pacing precision, or attention guidance scores.","supporting_citations":[],"review_version":1}