REVIEW 3 major objections 5 minor 19 references
PolyPresentation: A Multimodal AI Platform for Slide-Aware Iterative Presentation Practice
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Presentation coaching should be iterative and slide-aware, this paper argues: PolyPresentation links every feedback item to the slide and moment where the problem occurred, and its evidence-linked reports align with human ratings and…
desk verdict A slide-aware practice loop that is honestly built and honestly described, but the main feedback-quality comparison uses GPT-5 to judge a system that itself runs on GPT-5, so the headline superiority claim needs a blind human eval before it can be accepted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the slide-grounded evidence record. Every captured signal—speech transcript, slide-switch events, per-slide timing, keyword coverage, Q&A turns—is timestamped and aligned to the slide timeline, so each rubric judgment and recommendation can point to a slide number and a transcript span. On top of this record, the platform runs a rubric-based evaluator, an action-plan generator that decides whether a problem calls for delivery practice or slide revision, and a refinement agent that proposes deck edits while preserving the presenter's argument.
What would settle it
Have a panel of independent human raters, blind to source, score the same 20 feedback reports on the same seven dimensions with the same weights; if the platform no longer holds the top weighted score, the judged-superiority claim fails.
Extended reading notes
Core claim
On its own terms, the discovery is that slide-grounded evidence can serve as the organizing spine of presentation-practice feedback. Instead of scoring a finished delivery, PolyPresentation reconstructs what happened on each slide—what the presenter said, how long they stayed, which keywords were covered, whether a transition was rushed, how Q&A answers went—and uses that aligned record to produce rubric scores, an action plan, and deck revision suggestions. The claimed result is that this makes automated feedback more actionable, context-aware, and practice-oriented than the single-run feedback of existing systems, and the evaluation reports the highest overall feedback-quality score among five systems (7.54) with first-place rankings on 18 of 20 samples.
Load-bearing premise
The headline result assumes that the large language model acting as judge, which also powers several of the platform's modules, rates all systems' feedback without favoring output that resembles its own style; the paper concedes this comparison may be biased.
Editorial extensions
If this is right
- A presenter can close a full loop in one session: practice, rehearse, answer simulated questions, receive evidence-linked feedback, revise slides, and then practice again with explicit next-round goals.
- Feedback distinguishes delivery issues from coverage and deck-design issues, since each problem is tied to slide context rather than to isolated behavioral indicators.
- Unavailable or unreliable modalities are reported as not assessed instead of being guessed, making the feedback more honest about its own evidentiary basis.
- The largest reported gains over the closest coaching baseline occur in coverage, depth, and transfer, suggesting the loop helps with diagnosing missing content and planning future practice, not just polishing delivery.
- Q&A turns are treated as evidence of audience readiness, so practice extends beyond prepared speech to handling questions.
Reading between the lines
- A natural transfer target is interview practice, teaching demonstrations, or medical communication, where feedback also needs to cite a specific artifact and plan the next attempt; the same evidence-alignment spine may generalize.
- The evaluation does not test multi-round improvement, so the strongest form of the paper's thesis—that repeated loop use makes presentations objectively better—remains an untested extension; a longitudinal study measuring score gains across rounds would settle it.
- Because the judging model also powers several of the platform's modules, part of the reported superiority could reflect judge self-preference; a blinded human-rater replication is the direct check.
- The not-assessed policy implies the platform's usefulness depends on which modalities are available, so adding gaze or prosody traces could change rubric scores and action plans; the architecture is modular with respect to evidence richness.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces PolyPresentation, a multimodal AI platform for slide-aware iterative presentation practice. The system combines slide-by-slide practice, full rehearsal, simulated audience Q&A, evidence-grounded feedback, and slide deck refinement into a unified practice loop. The authors report two evaluations: rubric-scoring alignment with human ratings on 20 academic presentation rehearsals (Table 1) and a feedback-quality comparison against four baseline systems using a GPT-5 judge over a seven-dimension rubric (Table 2). The paper claims that PolyPresentation provides more actionable, context-aware, and practice-oriented support than existing systems, and it acknowledges in Section 6 that the feedback-quality comparison may be subject to evaluation bias because GPT-5 is used both as the judge and within PolyPresentation itself.
Significance. If the central claim were fully supported, PolyPresentation would be a useful contribution to AI-assisted presentation coaching: the slide-grounded evidence construction, the explicit link between feedback and practice planning, and the design choice to report unavailable modalities as 'not assessed' are all sensible and potentially valuable. The human-alignment analysis in Table 1, with pooled r=0.836, QWK=0.830, ICC(2,1)=0.831, and MAE=0.34, is moderately encouraging evidence that the platform's rubric scores resemble human judgments. Credit is also due to the authors for stating the main limitation of the GPT-5-based comparison explicitly in Section 6 rather than hiding it. However, the headline superiority claim rests on a non-independent evaluation, and the human-alignment results do not measure the actionability, context-awareness, or practice-orientation that the abstract emphasizes.
major comments (3)
- [Section 5.2 and Section 6] The feedback-quality comparison in Section 5.2, which is the only quantitative evidence for the claim that PolyPresentation provides more actionable, context-aware, and practice-oriented support, uses GPT-5 as the judge while PolyPresentation's own real-time hints, feedback generation, and Q&A modules also run on GPT-5 (Section 3.2). This is a self-referential evaluation: the judge and the judged system share the same underlying model, so the reported ranking in Table 2 could reflect GPT-5 favoring outputs that resemble its own style or structure rather than genuine quality. The authors concede in Section 6 that 'this comparison may be subject to evaluation bias,' but the paper still advances the comparative claim as a central result. A blinded human evaluation of feedback quality, or at minimum an independent judge model not used in any of the compared systems, is required before Table 2 can support the headline claim.
- [Table 2 and Section 5.2] No statistical testing is reported for the differences in Table 2. The statement that PolyPresentation 'ranks first on 18 of 20 samples' is presented without a paired test, and the overall-score gap (7.54 vs. 6.73 for PresentCoach) is not accompanied by confidence intervals, standard deviations, or per-order variation even though scores were averaged over three presentation orders. The authors should report per-sample and per-order variability and apply paired significance tests (e.g., Wilcoxon signed-rank or Friedman) to establish whether the observed advantages are robust rather than noise.
- [Section 5.1 and Table 1] The human-alignment analysis validates only five rubric scores (Appropriateness, Analysis, Persuasiveness, Clarity, Interaction); it does not evaluate whether PolyPresentation's feedback is more actionable, context-aware, or practice-oriented than that of the baselines. These latter constructs are the distinctive claims in the abstract, and no human rating of them is provided. Additionally, the rater information is incomplete: the number of human raters, their expertise, the rating instructions, and the inter-rater reliability among human raters are not reported, which is important for interpreting QWK and ICC(2,1) values. The 'pooled' statistics also combine ratings across all five criteria and all 20 presentations, which may inflate apparent agreement by pooling heterogeneous items; per-criterion and per-presentation breakdowns with appropriate error estimates should be provided.
minor comments (5)
- [Section 3.2] The paper refers to 'GPT-5' as the underlying model for hints and feedback, but does not specify the exact model version, API parameters, or prompt details, which limits reproducibility.
- [Section 5.2 and Table 2] The 'PresentCoach (reprod.)' baseline is described only as a reproduction; the authors should state how the reproduction was adapted to the same multimodal evidence and rubric, and whether the original authors were consulted for fidelity.
- [Table 2] The phrase 'ranks first on 18 of 20 samples' should specify how ties are handled, since three presentation orders per sample could produce tied or inconsistent rankings.
- [Section 5.2] The dimension weights (20%, 20%, 15%, 15%, 15%, 10%, 5%) are introduced without justification or sensitivity analysis; a brief rationale or a robustness check would strengthen the weighted overall score.
- [Section 5.1] The caption of Table 1 reports pooled statistics but does not define the pooling procedure precisely; the authors should state whether the pooled values are computed by combining all 100 paired ratings or by averaging per-criterion values.
Circularity Check
Feedback-quality superiority rests on a GPT-5 judge that also powers PolyPresentation's own modules; Table 1's human alignment validates rubric scores, not the actionability claim.
-
other
[Section 5.2 (Overall Feedback Quality, Table 2); acknowledged in Section 6 (Conclusion)]
"For each of the 20 samples, GPT-5 evaluated five anonymized feedback reports against the same multimodal evidence in three different orders. Scores were averaged across the three orders to reduce order bias. ... Because GPT-5 is also used in several PolyPresentation modules, this comparison may be subject to evaluation bias."
The central claim that PolyPresentation provides more actionable, context-aware, and practice-oriented support is supported by Table 2, where GPT-5 is the judge. Section 3.2 states that PolyPresentation's real-time hints use GPT-5, and Section 3.3's evidence-grounded feedback is produced by the same LLM/VLM pipeline. Thus the scored feedback reports are generated by the same model that rates them, so the judge and the system are not independent. The 'quality' scores can reflect GPT-5's self-preference in style, structure, and language rather than an external quality standard. The paper's own limitation statement concedes this bias.
full rationale
The paper is an empirical systems paper without a formal derivation chain, so most of its content is not circular: the human rubric-alignment study in Section 5.1 uses independent human raters, and the related-work self-citations (e.g., [16], [17]) are not load-bearing for the central claim. The one significant circularity is the feedback-quality comparison in Section 5.2: PolyPresentation's output is generated with GPT-5, and the evaluation of that output uses GPT-5 as the judge. This is not an independent test of feedback quality; it is a self-referential evaluation, as the authors explicitly acknowledge in Section 6. The human-alignment results give independent evidence that PolyPresentation's rubric scores agree with human ratings, but they do not validate the 'more actionable, context-aware, and practice-oriented' claim that the paper's abstract advertises. Because the headline superiority result therefore rests on a non-independent rater, and because no blinded human evaluation or independent-model validation of feedback quality is provided, the circularity score is 6: partial circularity in the main comparative claim, while the rest of the platform's reported evaluation is independent.
Assumptions & free parameters
free parameters (1)
- Feedback rubric dimension weights =
0.20, 0.20, 0.15, 0.15, 0.15, 0.10, 0.05
assumptions (3)
- domain assumption GPT-5 judge produces valid, unbiased feedback-quality ratings.
- domain assumption Human ratings are a reliable ground truth for presentation quality.
- domain assumption 20 academic conference rehearsals are representative of presentation practice.
Cite this review
Pith. "Pith review of PolyPresentation: A Multimodal AI Platform for Slide-Aware Iterative Presentation Practice." pith.science (2026). https://pith.science/paper/K7IIPNQG
@misc{pith2026260812857,
author = {Pith},
title = {Pith review of: PolyPresentation: A Multimodal AI Platform for Slide-Aware Iterative Presentation Practice},
year = {2026},
howpublished = {\url{https://pith.science/paper/K7IIPNQG}},
note = {Machine review of arXiv:2608.12857}
}
read the original abstract
Presentations are essential for students, researchers, and professionals to communicate ideas persuasively, yet delivering them effectively requires repeated practice that coordinates content, delivery, visual materials, and audience interaction. Existing AI-assisted rehearsal tools provide scalable feedback, but they often treat presentations as single-run delivery performances, offering limited support for linking feedback to the slide deck or planning what to practice in the next iteration. To address this gap, we introduce PolyPresentation, a multimodal AI platform for slide-aware iterative presentation practice. PolyPresentation organizes slide-by-slide practice, full rehearsal, audience Q&A, and feedback into a unified practice loop, using slide-grounded evidence to help presenters diagnose performance issues and prepare for subsequent practice. We evaluate PolyPresentation through a rubric-based comparison with four baseline systems on 20 academic presentation rehearsals, and additionally assess its alignment with human ratings. Results suggest that PolyPresentation provides more actionable, context-aware, and practice-oriented support for improving presentations. The demonstration video is available at https://youtu.be/MmWj9O_PJxw.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
In: Joint Proceedings of HEXED-L3MNGET 2024
Cha, J., Han, J., Yoo, H., Oh, A.: CHOP: Integrating ChatGPT into EFL Oral Presentation Practice. In: Joint Proceedings of HEXED-L3MNGET 2024. CEUR Workshop Proceedings, vol. 3840 (2024). https://ceur-ws.org/Vol-3840/ L3MNGET24_paper6.pdf
work page 2024
-
[2]
arXiv preprint arXiv:2511.15253 (2025)
Chen, S., Zhou, J., Xu, X., Yang, X., Guo, L., Chen, Y.-C.: PresentCoach: Dual- Agent Presentation Coaching through Exemplars and Interactive Feedback. arXiv preprint arXiv:2511.15253 (2025). https://doi.org/10.48550/arXiv.2511.15253
-
[3]
arXiv preprint arXiv:2603.07244 (2026)
Chen, X.-S., Zhu, J., Li, P.-l., Wang, H., Yang, S., Guo, M.-H.: PresentBench: A Fine-Grained Rubric-Based Benchmark for Slide Generation. arXiv preprint arXiv:2603.07244 (2026). https://doi.org/10.48550/arXiv.2603.07244
-
[4]
Chollet, M., Wörtwein, T., Morency, L.-P., Shapiro, A., Scherer, S.: Exploring Feedback Strategies to Improve Public Speaking: An Interactive Virtual Audience Framework. In: UbiComp 2015 (2015). https://doi.org/10.1145/2750858.2806060 10 C. Chen et al
arXiv 2015
-
[5]
Fu, T.-J., Wang, W.Y., McDuff, D., Song, Y.: DOC2PPT: Automatic Presentation Slides Generation from Scientific Documents. In: AAAI 2022, pp. 634–642 (2022). https://doi.org/10.1609/aaai.v36i1.19943
-
[6]
Learning and Individual Differences103, 102274 (2023)
Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., et al.: ChatGPT for Good? On Opportunities and Challenges of Large Language Models for Education. Learning and Individual Differences103, 102274 (2023). https://doi.org/10.1016/ j.lindif.2023.102274
arXiv 2023
-
[7]
Leong, C.W., Jawahar, N., Basheerabad, V., Wörtwein, T., Emerson, A., Sivan, G.: Combining Generative and Discriminative AI for High-Stakes Interview Practice. In: ICMI Companion 2024, pp. 94–96 (2024). https://doi.org/10.1145/3686215. 3688377
-
[8]
arXiv preprint arXiv:2512.04529 (2025)
Liang, X., Zhang, X., Xu, Y., Sun, S., You, C.: SlideGen: Collaborative Multimodal Agents for Scientific Slide Generation. arXiv preprint arXiv:2512.04529 (2025). https://doi.org/10.48550/arXiv.2512.04529
Show all 19 references
-
[9]
Journal of Learning Analytics11(3), 224–248 (2024)
Ochoa, X., Zhao, H.: OpenOPAF: An Open Source Multimodal System for Au- tomated Feedback for Oral Presentations. Journal of Learning Analytics11(3), 224–248 (2024). https://doi.org/10.18608/jla.2024.8411
2024
-
[10]
In: ICMI 2015, pp
Schneider, J., Börner, D., van Rosmalen, P., Specht, M.: Presentation Trainer, your Public Speaking Multimodal Coach. In: ICMI 2015, pp. 539–546 (2015). https: //doi.org/10.1145/2818346.2830603
2015
- [11]
-
[12]
IEEE Access11, 84013–84026 (2023)
Su, S.Y.T., Okada, S., Huang, H.-H., Leong, C.W.: Multimodal Transfer Learning for Oral Presentation Assessment. IEEE Access11, 84013–84026 (2023). https: //doi.org/10.1109/ACCESS.2023.3295832
2023
- [13]
-
[14]
In: NAACL-HLT 2021, pp
Sun, E., Hou, Y., Wang, D., Zhang, Y., Wang, N.X.R.: D2S: Document-to-Slide Generation Via Query-Based Text Summarization. In: NAACL-HLT 2021, pp. 1405–1418 (2021). https://doi.org/10.18653/v1/2021.naacl-main.111
2021 doi
-
[15]
In: IUI 2015, pp
Tanveer, M.I., Lin, E., Hoque, M.E.: Rhema: A Real-Time In-Situ Intelligent In- terface to Help People with Public Speaking. In: IUI 2015, pp. 286–295 (2015). https://doi.org/10.1145/2678025.2701386
2015
- [16]
-
[17]
arXiv preprint arXiv:2607.10310 (2026)
Wen, Z., Cao, J., Chan, K., Wang, Z., Chen, C., Liu, X., Yin, J., Li, Z.: Poly- Interview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment. arXiv preprint arXiv:2607.10310 (2026). https://doi.org/10.48550/arXiv.2607.10310
- [18]
-
[19]
In: EMNLP 2025, pp
Zheng, H., Guan, X., Kong, H., Zhang, W., Zheng, J., Zhou, W., Lin, H., Lu, Y., Han, X., Sun, L.: PPTAgent: Generating and Evaluating Presentations Be- yond Text-to-Slides. In: EMNLP 2025, pp. 14402–14418 (2025). https://doi.org/ 10.18653/v1/2025.emnlp-main.728
2025 doi
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.