{"id":"5cd9ce30-9f01-45d2-a337-a1d164d1e210","arxiv_id":"2608.12857","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"PolyPresentation, a slide-aware AI rehearsal platform, scores highest among five systems in a feedback-quality evaluation, though the comparison uses the same AI model as judge and system.","lead":"This paper presents PolyPresentation, an AI platform that guides presenters through repeated practice by attaching feedback to their slides and tracking progress across rounds. It is worth reading because it tests whether AI coaching can move beyond one-shot scoring toward genuinely iterative presentation improvement.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim of superior actionable, context-aware support rests on a GPT-5 judge that also powers the proposed system; Table 1's human alignment only validates rubric scores, not feedback actionability. A blinded human evaluation is needed.","rationale":"I agree with the reader that the weakest assumption is the impartiality of the GPT-5 judge in Section 5.2. This is the most load-bearing point because every other part of the evaluation either does not address the central claim or is not independent. The system's own modules are GPT-5-based (Section 3.2), and the authors admit the bias risk (Section 6). The human-alignment study (Table 1) is a useful positive control for rubric-scoring consistency, but it says nothing about whether the feedback helps presenters improve, which is the core of the claimed contribution. Even if the GPT-5 judge were unbiased, the absence of significance testing means the small sample (n=20) could yield unstable rankings; a single outlier or order effect could change the ordering. Thus the paper's headline claim is conditional: it holds only if an independent, blinded evaluation confirms the GPT-5-judged ranking. The proposed concrete test—a blinded human rater study with a paired significance test—directly addresses this. Such a study is feasible and is also listed by the authors as future work. Because the concern is real but potentially addressable, the reader's CONDITIONAL verdict is appropriate, and my read does not change it.","tokens_in":6909,"tokens_out":4786,"duration_ms":44629,"concrete_test":"Recruit three independent human raters with presentation-coaching or academic-presentation experience. For each of the 20 rehearsal samples, present the five anonymized system feedback reports (PolyPresentation, PresentCoach, VLM, LLM, Rule-based) in randomized order alongside the same multimodal evidence used in Sect. 5.2. Ask raters to score all seven rubric dimensions (Validity, Coverage, Impact, Depth, Actionability, Transfer, Organization) on the same 1-10 scale and compute the weighted overall score per system. Report the mean weighted scores and run a paired Wilcoxon signed-rank test (or mixed-effects model) comparing PolyPresentation with PresentCoach. If PolyPresentation no longer ranks first, or if the difference is not statistically significant (p < 0.05), the Table 2 result does not support the headline superiority claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Abstract) is that PolyPresentation provides 'more actionable, context-aware, and practice-oriented support' than existing systems. Table 2, the sole quantitative evidence for this claim, is produced by GPT-5 as judge over five systems' feedback reports. However, GPT-5 is the model underlying PolyPresentation's own real-time hints, feedback generation, and Q&A modules (Sect. 3.2). The authors explicitly concede 'this comparison may be subject to evaluation bias' (Sect. 6). The human-alignment analysis in Table 1, while positive, measures agreement on five rubric scores (Appropriateness, Analysis, Persuasiveness, Clarity, Interaction) between the platform and human raters; it does not measure whether the feedback is more actionable, context-aware, or practice-oriented. Thus the headline superiority claim is supported only by a judge that is not independent of the system being judged. In addition, no significance testing is reported for the Table 2 differences (e.g., the 7.54 vs. 6.73 weighted score gap), and the 'ranks first on 18 of 20 samples' statement lacks a paired test. Because the same GPT-5 judge may favor text that resembles its own output (e.g., style, structure, evidence citation), the observed ranking could reflect self-preference rather than genuine quality. This is the load-bearing weakness: the key outcome variable is measured by a non-independent rater, and no human or independent-model validation of feedback quality is provided.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces PolyPresentation, a multimodal AI platform for slide-aware iterative presentation practice. The system combines slide-by-slide practice, full rehearsal, simulated audience Q&A, evidence-grounded feedback, and slide deck refinement into a unified practice loop. The authors report two evaluations: rubric-scoring alignment with human ratings on 20 academic presentation rehearsals (Table 1) and a feedback-quality comparison against four baseline systems using a GPT-5 judge over a seven-dimension rubric (Table 2). The paper claims that PolyPresentation provides more actionable, context-aware, and practice-oriented support than existing systems, and it acknowledges in Section 6 that the feedback-quality comparison may be subject to evaluation bias because GPT-5 is used both as the judge and within PolyPresentation itself.","tokens_in":7253,"tokens_out":2489,"duration_ms":25786,"significance":"If the central claim were fully supported, PolyPresentation would be a useful contribution to AI-assisted presentation coaching: the slide-grounded evidence construction, the explicit link between feedback and practice planning, and the design choice to report unavailable modalities as 'not assessed' are all sensible and potentially valuable. The human-alignment analysis in Table 1, with pooled r=0.836, QWK=0.830, ICC(2,1)=0.831, and MAE=0.34, is moderately encouraging evidence that the platform's rubric scores resemble human judgments. Credit is also due to the authors for stating the main limitation of the GPT-5-based comparison explicitly in Section 6 rather than hiding it. However, the headline superiority claim rests on a non-independent evaluation, and the human-alignment results do not measure the actionability, context-awareness, or practice-orientation that the abstract emphasizes.","major_comments":[{"comment":"The feedback-quality comparison in Section 5.2, which is the only quantitative evidence for the claim that PolyPresentation provides more actionable, context-aware, and practice-oriented support, uses GPT-5 as the judge while PolyPresentation's own real-time hints, feedback generation, and Q&A modules also run on GPT-5 (Section 3.2). This is a self-referential evaluation: the judge and the judged system share the same underlying model, so the reported ranking in Table 2 could reflect GPT-5 favoring outputs that resemble its own style or structure rather than genuine quality. The authors concede in Section 6 that 'this comparison may be subject to evaluation bias,' but the paper still advances the comparative claim as a central result. A blinded human evaluation of feedback quality, or at minimum an independent judge model not used in any of the compared systems, is required before Table 2 can support the headline claim.","section":"Section 5.2 and Section 6"},{"comment":"No statistical testing is reported for the differences in Table 2. The statement that PolyPresentation 'ranks first on 18 of 20 samples' is presented without a paired test, and the overall-score gap (7.54 vs. 6.73 for PresentCoach) is not accompanied by confidence intervals, standard deviations, or per-order variation even though scores were averaged over three presentation orders. The authors should report per-sample and per-order variability and apply paired significance tests (e.g., Wilcoxon signed-rank or Friedman) to establish whether the observed advantages are robust rather than noise.","section":"Table 2 and Section 5.2"},{"comment":"The human-alignment analysis validates only five rubric scores (Appropriateness, Analysis, Persuasiveness, Clarity, Interaction); it does not evaluate whether PolyPresentation's feedback is more actionable, context-aware, or practice-oriented than that of the baselines. These latter constructs are the distinctive claims in the abstract, and no human rating of them is provided. Additionally, the rater information is incomplete: the number of human raters, their expertise, the rating instructions, and the inter-rater reliability among human raters are not reported, which is important for interpreting QWK and ICC(2,1) values. The 'pooled' statistics also combine ratings across all five criteria and all 20 presentations, which may inflate apparent agreement by pooling heterogeneous items; per-criterion and per-presentation breakdowns with appropriate error estimates should be provided.","section":"Section 5.1 and Table 1"}],"minor_comments":[{"comment":"The paper refers to 'GPT-5' as the underlying model for hints and feedback, but does not specify the exact model version, API parameters, or prompt details, which limits reproducibility.","section":"Section 3.2"},{"comment":"The 'PresentCoach (reprod.)' baseline is described only as a reproduction; the authors should state how the reproduction was adapted to the same multimodal evidence and rubric, and whether the original authors were consulted for fidelity.","section":"Section 5.2 and Table 2"},{"comment":"The phrase 'ranks first on 18 of 20 samples' should specify how ties are handled, since three presentation orders per sample could produce tied or inconsistent rankings.","section":"Table 2"},{"comment":"The dimension weights (20%, 20%, 15%, 15%, 15%, 10%, 5%) are introduced without justification or sensitivity analysis; a brief rationale or a robustness check would strengthen the weighted overall score.","section":"Section 5.2"},{"comment":"The caption of Table 1 reports pooled statistics but does not define the pooling procedure precisely; the authors should state whether the pooled values are computed by combining all 100 paired ratings or by averaging per-criterion values.","section":"Section 5.1"}],"recommendation":"major_revision","confidential_remarks":"The system contribution is interesting and the authors are transparent about the main limitation, but the comparative evaluation is not currently adequate for the strength of the claim. Because the flaw is localized to the feedback-quality evaluation and a blinded human study appears feasible within the paper's scope, major revision is more appropriate than rejection. I would also encourage the editor to ask for the raw scores and analysis scripts for Table 2 to enable verification of the ranking claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real systems contribution with an honest, central limitation that weakens its own headline. The platform is coherent: it links slide-level timing, transcripts, Q&A turns, and deck context into evidence-grounded feedback, then proposes deck revisions and next-round actions. Prior work mostly does one or the other — slide generation or rehearsal feedback — and I have not seen the two integrated in a single loop before. The writing is clear and the workflow is believable.\n\nWhat the paper does well: the human-alignment study in Section 5.1 is a genuine positive. Pooled r=0.836, QWK=0.830, ICC=0.831, MAE=0.34 across 20 academic presentations, with strong agreement on Analysis (r=0.910) and Appropriateness (r=0.860). That tells me the rubric scores are at least in the right neighborhood. The limitations section is also unusually candid: the authors explicitly say in Section 6 that using GPT-5 as judge while GPT-5 runs system modules creates evaluation bias.\n\nThe soft spots, in order:\n- The load-bearing problem is Table 2. That table is the only quantitative evidence for the claim that PolyPresentation provides more actionable, context-aware, and practice-oriented support. The judge is GPT-5, the same model that underlies the system's feedback and hints. This is not a minor caveat; it creates a real risk of self-preference. The authors' disclosure is honest but does not fix the inference.\n- The human-alignment analysis validates rubric scores, not feedback actionability. So it does not rescue Table 2.\n- The paper never tests the iterative loop. The claim is that repeated practice with slide-aware feedback improves presentations, but the evaluation is single-session output from one rehearsal.\n- Small sample, no significance testing, and no reporting of who the human raters were or how many of them rated.\n- The rubric dimension weights are set ad hoc, and the weighted overall score could shift if the weights changed.\n\nCitation pattern looks fine; the related work is mostly real and fairly cited.\n\nMy recommendation: send it to peer review. The system contribution is solid enough, and the limitation is named rather than hidden. It needs major revision, not rejection: a blinded human evaluation of feedback quality, or at least a judge independent of the system, plus significance tests and rater details. With that, it could be a useful addition to AI-assisted presentation coaching.","headline":"A slide-aware practice loop that is honestly built and honestly described, but the main feedback-quality comparison uses GPT-5 to judge a system that itself runs on GPT-5, so the headline superiority claim needs a blind human eval before it can be accepted.","tokens_in":7736,"tokens_out":3636,"would_cite":false,"duration_ms":35947,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Presentation coaching should be iterative and slide-aware, this paper argues: PolyPresentation links every feedback item to the slide and moment where the problem occurred, and its evidence-linked reports align with human ratings and…","keywords":["presentation practice","slide-aware feedback","multimodal AI","human-centered AI","iterative rehearsal","evidence-grounded feedback","rubric-based evaluation"],"falsifier":"Have a panel of independent human raters, blind to source, score the same 20 feedback reports on the same seven dimensions with the same weights; if the platform no longer holds the top weighted score, the judged-superiority claim fails.","tokens_in":6761,"feed_emoji":"🎤","tokens_out":8531,"duration_ms":81494,"temperature":0.7,"pith_summary":"The paper claims that automated presentation coaching works better when it treats rehearsal as an iterative, slide-aware loop rather than a one-shot performance. PolyPresentation, the platform it introduces, records slide-by-slide practice, full rehearsal, and simulated audience questions, then aligns the transcript, timing, slide switches, and Q&A answers along a slide timeline so that every piece of feedback can cite the exact slide and moment where a problem occurred. The authors report that its rubric scores agree with human ratings on 20 academic talk rehearsals (pooled Pearson correlation, weighted kappa, and agreement coefficient all above 0.83, with a mean absolute error of 0.34 on a 0–10 scale), and that in a large-language-model-judged comparison its feedback reports lead four baselines with the highest overall weighted score, 7.54 out of 10, on six of seven quality dimensions. The significance, if the results hold, is that presenters get concrete next-round actions and slide revisions tied to evidence instead of generic delivery advice.","feed_headline":"Practice feedback that cites the slide beats generic AI coaching","feed_subtitle":"The platform's slide-grounded loop outscores four baseline coaches in a 20-rehearsal comparison.","key_machinery":"The load-bearing mechanism is the slide-grounded evidence record. Every captured signal—speech transcript, slide-switch events, per-slide timing, keyword coverage, Q&A turns—is timestamped and aligned to the slide timeline, so each rubric judgment and recommendation can point to a slide number and a transcript span. On top of this record, the platform runs a rubric-based evaluator, an action-plan generator that decides whether a problem calls for delivery practice or slide revision, and a refinement agent that proposes deck edits while preserving the presenter's argument.","core_discovery":"On its own terms, the discovery is that slide-grounded evidence can serve as the organizing spine of presentation-practice feedback. Instead of scoring a finished delivery, PolyPresentation reconstructs what happened on each slide—what the presenter said, how long they stayed, which keywords were covered, whether a transition was rushed, how Q&A answers went—and uses that aligned record to produce rubric scores, an action plan, and deck revision suggestions. The claimed result is that this makes automated feedback more actionable, context-aware, and practice-oriented than the single-run feedback of existing systems, and the evaluation reports the highest overall feedback-quality score among five systems (7.54) with first-place rankings on 18 of 20 samples.","pith_inferences":["A natural transfer target is interview practice, teaching demonstrations, or medical communication, where feedback also needs to cite a specific artifact and plan the next attempt; the same evidence-alignment spine may generalize.","The evaluation does not test multi-round improvement, so the strongest form of the paper's thesis—that repeated loop use makes presentations objectively better—remains an untested extension; a longitudinal study measuring score gains across rounds would settle it.","Because the judging model also powers several of the platform's modules, part of the reported superiority could reflect judge self-preference; a blinded human-rater replication is the direct check.","The not-assessed policy implies the platform's usefulness depends on which modalities are available, so adding gaze or prosody traces could change rubric scores and action plans; the architecture is modular with respect to evidence richness."],"forward_implications":["A presenter can close a full loop in one session: practice, rehearse, answer simulated questions, receive evidence-linked feedback, revise slides, and then practice again with explicit next-round goals.","Feedback distinguishes delivery issues from coverage and deck-design issues, since each problem is tied to slide context rather than to isolated behavioral indicators.","Unavailable or unreliable modalities are reported as not assessed instead of being guessed, making the feedback more honest about its own evidentiary basis.","The largest reported gains over the closest coaching baseline occur in coverage, depth, and transfer, suggesting the loop helps with diagnosing missing content and planning future practice, not just polishing delivery.","Q&A turns are treated as evidence of audience readiness, so practice extends beyond prepared speech to handling questions."],"supporting_citations":[{"why":"Supplies the dual-agent coaching system that serves as the closest reproduced baseline in the feedback-quality comparison.","marker":"[2]"},{"why":"Supplies the oral-presentation practice framework whose prompt design underlies the LLM baseline.","marker":"[1]"},{"why":"Represents the real-time delivery coach that motivates shifting from isolated behavior cues to slide-grounded evidence.","marker":"[15]"},{"why":"Represents the multimodal public-speaking coach that the paper argues lacks slide context.","marker":"[10]"},{"why":"Provides the interactive virtual-audience framework that motivates the simulated audience and Q&A stage.","marker":"[4]"},{"why":"Represents the open multimodal oral-feedback platform that the slide-alignment approach extends.","marker":"[9]"},{"why":"Illustrates the separate slide-generation line of work that the paper contrasts with rehearsal evaluation.","marker":"[5]"},{"why":"Marshals the discussion of LLM reliability in education that motivates evidence grounding and the not-assessed policy.","marker":"[6]"},{"why":"Provides the closest slide-aware system that synchronizes narration with slides but lacks an iterative feedback loop.","marker":"[11]"}],"fun_headline_variants":["Slide-aware feedback loops outscore generic coaching in 20 runs","PolyPresentation: AI rehearsals pin feedback to your slides","Multimodal AI ties practice feedback to the slide deck","Slide-grounded AI coach beats baselines in presentation practice","Iterative practice AI uses slide evidence to improve feedback"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline result assumes that the large language model acting as judge, which also powers several of the platform's modules, rates all systems' feedback without favoring output that resembles its own style; the paper concedes this comparison may be biased.","fun_headline_variants_meta":{"raw":{"variants":["Slide-aware feedback loops outscore generic coaching in 20 runs","PolyPresentation: AI rehearsals pin feedback to your slides","Multimodal AI ties practice feedback to the slide deck","Slide-grounded AI coach beats baselines in presentation practice","Iterative practice AI uses slide evidence to improve feedback"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000155,"raw_usage":{"total_tokens":1178,"prompt_tokens":871,"completion_tokens":307,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":487,"completion_tokens_details":{"reasoning_tokens":240}},"tokens_in":487,"tokens_out":307,"duration_ms":3161,"temperature":1.0,"reasoning_tokens":240,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:51:46.901723+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a panel of independent human raters, blind to source, score the same 20 feedback reports on the same seven dimensions with the same weights; if the platform no longer holds the top weighted score, the judged-superiority claim fails.","supporting_citations":[{"cited_title":"arXiv preprint arXiv:2511.15253 (2025)","cited_arxiv_id":null,"evidence_quote":"Supplies the dual-agent coaching system that serves as the closest reproduced baseline in the feedback-quality comparison."},{"cited_title":"In: Joint Proceedings of HEXED-L3MNGET 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the oral-presentation practice framework whose prompt design underlies the LLM baseline."},{"cited_title":"In: ICMI 2015, pp","cited_arxiv_id":null,"evidence_quote":"Represents the multimodal public-speaking coach that the paper argues lacks slide context."},{"cited_title":"PresentAgent: Multimodal Agent for Presentation Video Generation","cited_arxiv_id":"2507.04036","evidence_quote":"Provides the closest slide-aware system that synchronizes narration with slides but lacks an iterative feedback loop."}],"review_version":1}