{"id":"8f9e3c22-84d4-4090-84a6-ac2ba82cd103","arxiv_id":"2608.08535","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An LLM-based watch-practice-evaluate system for teaching videos increases early-stage teachers' engagement and their ability to transfer observed instructional strategies to a new lesson.","lead":"TeachUp is an interactive system that helps early-stage teachers learn teaching strategies from classroom videos by detecting strategy moments, prompting reflection while watching, and offering guided microteaching practice with AI feedback. In a controlled study with 16 teachers, it raised engagement and improved how well they applied the learned strategies to a new teaching task, compared to watching videos and practicing on their own.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Primary learning-outcome result rests on single-rater holistic scores with no reliability evidence, and the paired comparison crosses two different raters and tasks.","rationale":"The reader identified the absence of inter-rater reliability as the weakest assumption, and I agree that this is the most load-bearing issue. The paper's central claim depends on a single pair of expert raters, each scoring only one subject area, with no evidence that the holistic Likert ratings are reproducible. The additional twist is that the within-subjects paired design crosses raters and tasks: each participant's TeachUp score comes from one rater and their baseline score from another, on different subject content. Z-score standardization does not fix this; it only adjusts the marginal distribution per rater. A concrete re-scoring study with fully crossed raters would directly test whether the effect is an artifact of measurement noise or rater expectations. The counterbalanced design, blind ratings, and controlled comparison are real strengths, and the qualitative results and interview study provide convergent support, so I do not recommend rejection. However, because the strongest quantitative result rests on an unvalidated single-rater measure, the paper should remain conditional until reliability evidence is supplied. This does not change the reader's verdict, but it sharpens the reason for conditionality.","tokens_in":21610,"tokens_out":4773,"duration_ms":58261,"concrete_test":"Recruit a second independent rater for each subject area, or have both original raters score all 32 teaching-test videos under fully crossed, condition-blinded conditions using the Section 5.2.4 rubric. Compute inter-rater reliability (weighted kappa or ICC) for Dim2; then re-run the TeachUp-vs-baseline comparison on the averaged ratings with a paired model that includes rater and task as factors. If ICC is below 0.6 or the Dim2 effect loses significance or has a wide confidence interval, the central claim is not supported by the current measurement. If reliability is high and the effect persists, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—TeachUp improves application of learned instructional strategies—is carried by Dim2 in Table 1 (p=.003). Section 5.2.4 reports that one mathematics teacher rated math teaching tests and one Chinese teacher rated Chinese teaching tests; no inter-rater reliability, rubric validation, or duplicate scoring is reported. Because each participant contributes one score from each rater on different subject tasks (Math with Cues/Advance Organizers vs. Chinese with Cooperative Learning), the paired TeachUp-vs-baseline difference is a contrast across two different raters and two different task contents. The Z-score standardization in Section 6 only normalizes each rater's marginal distribution; it cannot remove rater-by-condition or rater-by-task interactions. A holistic 7-point judgment of 'application' without demonstrated reliability could be influenced by rater expectations, perceived fluency, or task-difficulty differences, any of which could produce or mask the p=.003 effect. This is load-bearing because Dim2 is the primary RQ1 result; the other performance dimensions are marginal or nonsignificant (Dim1 p=.065, Dim3 p=.516, Dim4 p=.086). Section 8.4 does not address rater reliability, and the paper's own limitation list omits this threat to the main outcome measure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents TeachUp, an interactive system that helps early-stage teachers learn instructional strategies from classroom videos through a three-stage 'watch-practice-evaluate' loop. The authors first report a formative study (N=9) that motivates design requirements, then describe an LLM-powered pipeline that detects nine Marzano instructional strategies in classroom videos (claimed precision 0.634 at IoU=0.5), followed by a within-subjects study (N=16) comparing TeachUp with a baseline video-watching and self-practicing system. The main quantitative claims are that TeachUp improves participants' application of learned strategies to a new teaching task (Z-scored expert rating, p=.003), increases engagement (Q4, p=.026), confidence (Q1, p=.009), perceived practice benefit (Q3, p=.014), and intention to use (Q6, p=.017), while other performance dimensions are marginal or nonsignificant. Interviews with four in-service teachers are used to generalize the findings.","tokens_in":21854,"tokens_out":2508,"duration_ms":30297,"significance":"If the central effectiveness claims hold, TeachUp would be a useful contribution to HCI and teacher professional development, since it directly addresses a known gap in video-based learning: helping novice teachers notice, rehearse, and reflect on context-dependent instructional strategies. The paper also contributes a released benchmark of 89 annotated strategy segments and a reproducible LLM-based detection pipeline, which are valuable starting points even with modest precision. The study design is thoughtful in several respects: the user study uses a within-subjects design with counterbalanced tasks, blind expert raters, a pre-experiment preparation phase, and an incentive structure, and the qualitative data are analyzed with a transparent thematic approach. The paper is also honest about several limitations (e.g., the absence of a component-level ablation and the lack of long-term behavioral outcomes). However, the primary learning-outcome measure rests on single-rater holistic judgments with no reliability evidence, and the pipeline evaluation is an in-sample fit, so the strongest empirical claims are not yet fully supported.","major_comments":[{"comment":"The central RQ1 result—that TeachUp improves application of learned instructional strategies (p=.003)—rests entirely on holistic 7-point Likert scores given by a single expert rater per participant, with no inter-rater reliability, rubric validation, or duplicate scoring. Section 5.2.4 states that one mathematics teacher rated math tests and one Chinese teacher rated Chinese tests; each participant therefore contributes one score from each of two different raters on two different subject tasks (Cues/Advance Organizers in math vs. Cooperative Learning in Chinese). The Z-score standardization in Section 6 normalizes only each rater's marginal distribution and cannot remove rater-by-condition or rater-by-task interactions. A holistic judgment of 'application' with no demonstrated reliability could be influenced by rater expectations, perceived fluency, or task-difficulty differences, any of which could produce or mask the reported effect. This is load-bearing because Dim2 is the primary performance claim, while Dim1, Dim3, and Dim4 are marginal or nonsignificant. I ask the authors to provide inter-rater reliability on at least a subset of double-scored test videos, a rubric that maps the four metrics to observable behaviors, and an analysis or explicit acknowledgment of the rater/task confound in the paired comparison (Section 8.4 does not currently address this).","section":"§5.2.4, Table 1 (Dim2)"},{"comment":"The claimed pipeline precision of 0.634 is measured on the same benchmark from which the few-shot examples were drawn. The text states that 'we incorporated transcript segments as positive few-shot examples and frequently misclassified segments as negative few-shot examples' and that these examples were designed using the benchmark videos; the precision is then reported on that same benchmark. This is an in-sample fit, not an out-of-sample estimate, so the reported precision is likely optimistic relative to what users would encounter on new classroom videos. The zero-shot comparison (precision=0.095) is useful but does not resolve the circularity. Please report a held-out or cross-validated evaluation, or clearly frame the reported number as a development-set fit. This matters because the pipeline is presented as an enabling contribution and because Section 7's interview claims about clip quality are based on pipeline outputs that were not independently error-checked.","section":"§4.2.1 (Pipeline technical evaluation)"},{"comment":"The paper reports p-values for four expert-rated dimensions and six questionnaire items without any correction for multiple comparisons. Several effects are marginal (Dim1 p=.065, Dim4 p=.086, Q5 p=.053), and the headline effects (Dim2 p=.003, Q1 p=.009, Q3 p=.014, Q4 p=.026, Q6 p=.017) are drawn from the same set of participants across two conditions. Since the primary claims are selective, I ask the authors to report the number of comparisons considered, to apply or justify not applying a correction (e.g., Bonferroni or FDR), and to interpret marginal results accordingly.","section":"§6, Table 1"}],"minor_comments":[{"comment":"The arithmetic of the annotation agreement is unclear: the text reports 85 agreed events, 1 event added by one annotator, 4 events added by the other, and 5 disagreements resolved through discussion, which sums to 95 rather than the stated final benchmark of 89 instances. Please reconcile these counts.","section":"§4.2.1"},{"comment":"The sentence 'traditional inter-rater statistics such as Cohen's κ or ICC are not applicable' is too strong; segment-level agreement measures can be computed for temporal detection tasks, and the authors' IoU-based agreement already provides a reasonable alternative. A brief justification or reference would help.","section":"§4.2.1"},{"comment":"The two learning tasks differ in both subject and strategy, and the two tasks are assigned to conditions with counterbalancing only of task order and system order, not of task-condition pairing. Because each participant always does one task in one condition, any subject-level or strategy-level difficulty difference is fully confounded with condition in the paired analysis. This should at least be acknowledged in Section 8.4.","section":"§5.2.2"},{"comment":"Questionnaire items Q1–Q6 are reported with p-values but no per-item effect sizes or confidence intervals; reporting these would help readers assess practical significance alongside the Wilcoxon tests.","section":"Appendix C"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid design process and a credible prototype, and the authors' move to release the benchmark and prompts is commendable. My main concern is that the primary learning-outcome claim is carried by a single-rater holistic measure with no reliability evidence, and the paired design crosses raters and tasks. This is fixable within the manuscript's scope by adding duplicate scoring/IRR, a more detailed rubric, and a transparent discussion of the confound, plus a held-out evaluation of the detection pipeline. I do not think the results are so flawed that rejection is warranted, but the current version overstates the strength of the evidence for the headline RQ1 result."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, you should know this paper is a solid systems contribution for UIST. It builds TeachUp, a tool that turns recorded classroom videos into structured reflective practice for early-career teachers, and introduces a new benchmark for instructional strategy detection in that domain. The three-stage loop—watch with LAP-style hints, practice on generated microteaching scenarios, then compare your own video against the expert clip—is new and sensibly grounded in a formative study and in prior pedagogy (Marzano, LAP, Kolb). The authors also did a clean counterbalanced within-subjects study with blind expert raters, and the qualitative data is genuinely useful.\n\nThe main soft spot is the one the stress-test flagged: the headline learning outcome (Dim2, strategy application, p=.003) rests on a single expert rater per participant, with no inter-rater reliability or rubric validation. Since each participant is scored by the math teacher in one condition and the Chinese teacher in the other, condition is confounded with rater and task content. Z-score standardization removes a rater's overall strictness but not rater-by-condition or rater-by-task interactions. This matters because the other performance dimensions are marginal or null (Dim1 p=.065, Dim3 p=.516, Dim4 p=.086). The paper's own limitations list omits this threat.\n\nA smaller but real issue: the pipeline's 63.4% precision is in-sample—few-shot examples were chosen from the same benchmark used for evaluation—so the number is an optimistic feasibility estimate, not an independent performance measure. The authors do call it an 'initial feasibility check,' which is fair, but readers will likely quote it as a fact.\n\nNeither issue is fatal. The system is well-designed, the interviews with in-service teachers add external validity, and the authors are transparent about other limitations (participants are not real classroom teachers, no long-term outcomes, baseline is not an ablation). I would send this to peer review. A serious referee should require a second rater or an IRR report, and ideally a re-analysis with rater as a random effect, before letting the p=.003 claim stand. The pipeline section should also either present a held-out split or clearly label the precision as in-sample. Worth a reading-group discussion—it's a good case study in what counts as evidence in HCI evaluation.","headline":"A worthwhile systems and benchmark paper whose main learning-effect claim rests on single-rater scores with no reliability evidence; good enough for peer review, but treat the p=.003 as provisional.","tokens_in":22372,"tokens_out":6176,"would_cite":true,"duration_ms":61001,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"TeachUp, a system that structures video learning of teaching strategies into watch, practice, and evaluation stages, helps early-stage teachers apply what they learn from classroom videos to new lessons, outperforming a conventional…","keywords":["video-based learning","instructional strategies","reflective support","microteaching","LLM","classroom video analysis","teacher professional development","within-subjects study"],"falsifier":"Score the same set of 16 teaching-test videos with at least one additional independent rater blind to condition, using the same rubric, and compute inter-rater agreement; if agreement is low (intraclass correlation below roughly 0.6) or the TeachUp-versus-baseline difference on strategy application loses significance when rater identity is included in the model, the central claim would be undercut.","tokens_in":21427,"feed_emoji":"🎓","tokens_out":8654,"duration_ms":79819,"temperature":0.7,"pith_summary":"The paper argues that early-stage teachers learn instructional strategies from recorded classroom videos better when watching is wrapped in a structured reflective loop than when they watch and practice on their own. It presents TeachUp, which detects nine instructional-strategy categories in videos, prompts reflective questions while the video plays, generates a scenario-based microteaching task, and gives multimodal feedback comparing the user's performance with the original lesson. In a within-subjects study with 16 pre-service teachers, TeachUp users scored significantly higher on applying the learned strategy to a new teaching task (Z +0.32 vs -0.32, p = .003), and reported higher engagement, confidence, and intention to use the system than with the baseline. Interviews with four in-service teachers corroborate the usefulness of the strategy-clip detection and the reflective cycle, while also pointing to concerns about task flexibility and transfer to real classrooms.","feed_headline":"Reflective video loop lifts teachers' use of learned strategies","feed_subtitle":"In a 16-teacher test, TeachUp users scored +0.32 vs -0.32 on applying a strategy to a new lesson (p=.003).","key_machinery":"The central mechanism is the watch-practice-evaluate interaction loop, instantiated as a concrete implementation of an experiential learning cycle. The loop is powered by two computational components: a transcript-based detection pipeline that uses an ensemble of three large language models to segment classroom videos and tag clips with nine instructional-strategy categories, and a multimodal large language model that evaluates users' microteaching recordings and generates comparison-based feedback. The detection pipeline's rule-based aggregator merges overlapping predictions across models, and its ensemble strategy (accepting only clusters predicted by at least two models) is what yields the reported precision of 0.634.","core_discovery":"The core claim is that the implicit pedagogical reasoning behind a teacher's classroom actions can be made learnable by pairing video examples with a reflection-and-practice loop. TeachUp operationalizes this as a watch-practice-evaluate cycle: an LLM-powered pipeline segments classroom videos into clips tagged with one of nine instructional strategies; reflective hints derived from a claim-evidence-reasoning-alternative protocol guide attention while watching; a scenario generator creates a microteaching task in which the user rehearses the strategy; and a multimodal LLM evaluates the practice video against the original instructional case. The experimental evidence compares TeachUp with a baseline that offers the same videos and free self-practice, and finds the reflective loop significantly improves expert-rated application of the strategy to a brand-new lesson, along with engagement and confidence. The paper presents the strategy-detection pipeline (precision 0.634 at IoU = 0.5) as an enabling component and releases a benchmark of 89 annotated strategy instances as a starting point for future work.","pith_inferences":["If the headline effect replicates under multi-rater scoring, the watch-practice-evaluate loop could be embedded into existing teacher-education platforms as a lightweight asynchronous layer, since the pipeline is designed as plug-and-play middleware.","The single-rater outcome measure leaves open the possibility that the reported effect size is inflated by scoring noise; a cheap replication with two independent raters and a pre-registered rubric is the natural next step.","Because the detection pipeline works from transcripts alone, strategies that are primarily non-verbal (e.g., nonlinguistic representations, cooperative-learning arrangements) may be under-detected, which would bias which strategies learners get offered; this is a testable implication the paper does not address.","The transfer-distance tension that experienced teachers raised suggests an adaptive version of the practice generator: adjust how far the simulated scenario departs from the observed video based on the user's demonstrated understanding, something the current system does not attempt."],"forward_implications":["Early-stage teachers can transfer a strategy observed in a video to a new lesson more successfully when reflection and rehearsal are structured around the video, which suggests video libraries for teacher professional development should pair footage with interactive tasks rather than rely on self-directed viewing.","The transcript-based, LLM-driven detection pipeline offers a scalable way to index large collections of recorded classroom videos by instructional strategy, lowering the effort of finding relevant examples.","The significant gains in engagement and confidence imply that the reflective loop may reduce drop-off in self-paced online teacher learning, where motivation is a known barrier.","Because imitation of surface implementations showed no significant improvement, the effective mechanism appears to be understanding the strategy's timing and rationale rather than copying the observed teacher's actions.","The released benchmark of 89 annotated strategy instances provides a concrete test set for future work on automatic instructional-strategy detection in classroom videos."],"supporting_citations":[{"why":"Supplies the nine instructional-strategy categories that TeachUp detects, teaches, and uses to tag video clips.","marker":"[38]"},{"why":"Provides the Lesson Analysis Protocol (claim-evidence-reasoning-alternative) on which TeachUp's reflective hints during watching are based.","marker":"[54]"},{"why":"Defines the microteaching cycle (plan, teach, feedback) that TeachUp adapts for its practice stage.","marker":"[1]"},{"why":"Provides the experiential learning cycle that the watch-practice-evaluate loop is designed to instantiate.","marker":"[29]"},{"why":"Documents the barriers in video-based teacher learning (complexity, volume, missing reflection) that motivate the system's design.","marker":"[17]"},{"why":"Supplies the multimodal LLM used to evaluate users' microteaching videos and generate the feedback reports.","marker":"[61]"},{"why":"One of the three language models ensembled in the strategy-detection pipeline to improve precision.","marker":"[31]"},{"why":"Provides the coding method for verbal reflection responses, used to count valid inspiration points in the study.","marker":"[60]"}],"fun_headline_variants":["Reflective video loop boosts teacher strategy transfer","LLM-powered reflection improves video learning for teachers","Watch, reflect, practice: AI video coach for teachers","Video reflection sharpens new teachers' strategy use","AI-assisted video review lifts teaching performance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The main learning-outcome result rests on a single expert rater's blind 7-point Likert scoring of each participant's short teaching-test video as a valid and reliable measure of applying learned teaching strategies, and the paper reports no inter-rater reliability or rubric-validation evidence for this measure.","fun_headline_variants_meta":{"raw":{"variants":["Reflective video loop boosts teacher strategy transfer","LLM-powered reflection improves video learning for teachers","Watch, reflect, practice: AI video coach for teachers","Video reflection sharpens new teachers' strategy use","AI-assisted video review lifts teaching performance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000365,"raw_usage":{"total_tokens":1960,"prompt_tokens":933,"completion_tokens":1027,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":957}},"tokens_in":549,"tokens_out":1027,"duration_ms":11092,"temperature":1.0,"reasoning_tokens":957,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:32:30.313369+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Score the same set of 16 teaching-test videos with at least one additional independent rater blind to condition, using the same rubric, and compute inter-rater agreement; if agreement is low (intraclass correlation below roughly 0.6) or the TeachUp-versus-baseline difference on strategy application loses significance when rater identity is included in the model, the central claim would be undercut.","supporting_citations":[{"cited_title":"2001.Classroom instruction that works: Research-based strategies for increasing student achievement","cited_arxiv_id":null,"evidence_quote":"Supplies the nine instructional-strategy categories that TeachUp detects, teaches, and uses to tag video clips."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Lesson Analysis Protocol (claim-evidence-reasoning-alternative) on which TeachUp's reflective hints during watching are based."},{"cited_title":"Allen and K","cited_arxiv_id":null,"evidence_quote":"Defines the microteaching cycle (plan, teach, feedback) that TeachUp adapts for its practice stage."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the barriers in video-based teacher learning (complexity, volume, missing reflection) that motivate the system's design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the coding method for verbal reflection responses, used to count valid inspiration points in the study."}],"review_version":1}