{"id":"fa76ec55-6d59-4263-9e31-8dd3970c4691","arxiv_id":"2505.07161","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Applying off-the-shelf dialogue act and discourse relation models to math classroom and tutoring transcripts shows that non-talk-move utterances play functional roles in discourse, and are not just fillers.","lead":"This paper combines three layers of dialogue analysis, talk moves, dialogue acts, and discourse relations, on math teaching and tutoring datasets. It finds that utterances without talk moves, often ignored, carry important functions in guiding and structuring classroom dialogue.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim about T-NONE functions rests on unvalidated cross-domain DA/DR labels; the distributions in Fig. 7 and Table 3 are only as trustworthy as the Switchboard/Minecraft-trained models.","rationale":"The reader's weakest assumption identifies exactly the load-bearing issue: the dialogue act and discourse relation layers, which drive the novel findings about non-talk-move utterances, are produced by models trained on non-educational corpora and never validated on the target datasets. My analysis agrees with the CONDITIONAL verdict: the paper is transparent about this limitation, the human talk move annotations provide a solid base for the unigram and transition analyses, and the qualitative examples are suggestive. However, the quantitative claims that T-NONE utterances serve guiding, acknowledging, and structuring functions depend on model labels whose accuracy in this domain is unknown. A human agreement study would settle whether the concern lands. If validation shows acceptable agreement, the central claim stands; if not, the authors would need to either fine-tune the models, restrict conclusions to human-verified subsamples, or substantially weaken the claims about the role of T-NONE. Thus the existing CONDITIONAL verdict is appropriate and should remain unchanged.","tokens_in":21786,"tokens_out":2315,"duration_ms":26600,"concrete_test":"Stratified validation study: sample 200 T-NONE utterances from TalkMoves and 200 from SAGA22, balanced by speaker role and session. Have two annotators experienced in SWBD-DAMSL and SDRT independently label each utterance for dialogue act and discourse relation, then compute model-human agreement (accuracy, Cohen's kappa, and per-category F1) for the DA classifier and Llamipa. If agreement on the critical categories—Acknowledgment-(Backchannel), Action-directive, Statement-non-opinion, Continuation, Elaboration—falls substantially below the models' reported performance on Switchboard/MSDC, or if the top-3 DA/DR distributions from human labels differ by more than 10 percentage points from Figure 7/Table 3, the central claim should be re-evaluated using either in-domain fine-tuned models or human-verified labels on a restricted subset.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's most important claim—that utterances without talk moves are not mere fillers but serve guiding, acknowledging, and structuring functions—is supported mainly by model-generated dialogue act and discourse relation labels. Section 2.2 selects a DA classifier trained on SWBD-DAMSL (82.4 accuracy on Switchboard) and Llamipa, an SDRT parser fine-tuned on the Minecraft Structured Dialogue Corpus. Section 4.2 explicitly acknowledges that neither model was fine-tuned on TalkMoves or SAGA22, yet Sections 3.1.2 and 3.3.2 treat the resulting labels as reliable evidence: Figure 7 reports that T-NONE and S-NONE are dominated by Statement-non-opinion, Acknowledgment-(Backchannel), and Action-directive, and Table 3 lists percentages such as ContS. (43.74%) for T-RESTAT → T-NONE based on these labels. If these cross-domain models systematically mislabel educational discourse—for example, if classroom 'okay' utterances are tagged as Statement-non-opinion rather than Acknowledgment, or if tutorial directions are assigned different SDRT relations than Minecraft build instructions—then the claim that T-NONE guides, acknowledges, and structures discourse loses its quantitative foundation. The paper provides no human validation, no error analysis, and no confidence filtering on either model's outputs. This is not merely an 'outside consensus' issue; it is a correctness risk internal to the argument, because the central conclusion is inferred directly from these unvalidated labels. The human talk move annotations and the qualitative examples in Figures 15–17 give partial independent support, but they do not establish the distributional percentages or the significance attributed to specific DA/DR patterns.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a multi-perspective discourse analysis framework that combines domain-specific talk moves (human-annotated in the TalkMoves and SAGA22 datasets), automatic dialogue acts (SWBD-DAMSL-style, 43 tags), and automatic discourse relations (SDRT, 16 relations) to study mathematics teaching and tutoring dialogues. The authors perform unigram distributions, sequential talk-move transition analyses, and multi-view deep dives, and report that utterances without talk moves (T-NONE/S-NONE) are not mere fillers but contribute to guiding, acknowledging, and structuring classroom discourse. The stated goal is to provide actionable feedback for educators and design principles for AI agents in mathematics education.","tokens_in":22208,"tokens_out":4337,"duration_ms":44289,"significance":"The paper addresses a genuine and timely problem: most automated classroom-discourse feedback focuses only on theory-driven talk moves and ignores the majority of utterances that do not fit those moves. The integration of human-annotated talk moves with off-the-shelf dialogue act and discourse relation models over two large educational datasets is novel, and the open-sourced models and reproducible analysis pipeline are strengths. If the central claim about the role of non-talk-move utterances were supported by validated labels, the framework would be a useful step toward richer, action-oriented feedback and toward more capable AI tutoring agents. However, as presented, the main conclusion rests on automatic labels that have not been shown to be reliable on educational discourse, so the significance is conditional on that validation being supplied.","major_comments":[{"comment":"The paper's central claim—that T-NONE utterances are not fillers but 'guide, acknowledge, and structure' discourse (abstract; §4.1)—is inferred directly from dialogue act and discourse relation distributions produced by models trained on Switchboard (the DA model, §2.2) and Minecraft dialogue (Llamipa, §2.2), neither of which was validated on TalkMoves or SAGA22. The limitation is acknowledged in §4.2, but the empirical results in Fig. 7 and Table 3 are nevertheless presented as evidence for the central claim. Without a human-annotated evaluation sample, an error analysis, or confidence-based filtering that demonstrates these cross-domain labels are accurate on mathematics classroom and tutoring transcripts, the quantitative foundation of the main conclusion is missing. The authors should either provide such validation or substantially soften the claims and mark them as hypotheses to be tested with in-domain labels.","section":"§3.1.2, §3.3.2, Fig. 7, Table 3"},{"comment":"The sequential analysis depends on several ad-hoc thresholds: the 10% transition probability cutoff in Figs. 8 and 9, the 5% frequency exclusion for T-NONE counts in §2.3.2, and the top-3/top-7 DA selection in §2.3.1. No sensitivity analysis is provided, so it is unclear whether the reported patterns, such as the 'higher T-NONE interactions in tutoring' (§3.2.2), are robust to reasonable threshold changes. Additionally, Eq. (1) is under-specified: it appears to compute an expected number of intervening T-NONE utterances, but the treatment of zero T-NONE cases and the exclusion rule 'frequency below 5%' are not clearly defined. The authors should justify the thresholds or test several values.","section":"§2.3.2, Eq. (1), Figs. 8–11"},{"comment":"All cross-domain comparisons (teaching vs. tutoring) are descriptive, yet the text makes claims such as 'significantly higher' for the T-PRSREA→S-PROEVI transition (41% vs. 28%, §3.2.1) and states differences in T-NONE prevalence (5.8% and 4.3%) as meaningful. The data have a nested structure (utterances within sessions) and the two datasets differ in session count, number of students per session, and session length, which may confound simple frequency comparisons. The paper should report confidence intervals, effect sizes, or a multilevel model; at minimum, the wording should be 'descriptively different' rather than 'significant'.","section":"§3.2.1, §3.1"},{"comment":"The unigram DA analysis explicitly excludes the Continued-by-same-speaker dialogue act (§2.3.1), yet Table 3 reports ContS. as the most frequent DA in several T-NONE transition contexts (e.g., 43.74% for T-RESTAT→T-NONE and 61.66% for S-MCLAIM→T-NONE). This inconsistency between the bottom-up DA analysis and the top-down transition analysis is confusing and should be resolved: either apply the same exclusion/inclusion criterion everywhere or explain why ContS. is meaningful in the transition analysis but excluded from unigram distributions.","section":"§2.3.1, Table 3"}],"minor_comments":[{"comment":"The percentages in Figure 2 appear inconsistent with the text: the caption/legend lists T-None at 49.3% but the pie label shows 49.7%, and S-None is listed as 10.92% but labeled 11.0%; please reconcile.","section":"Figure 2"},{"comment":"The phrase 'flattened multi-functional SWBD-MASL schema' in the abstract and §1.2 is likely a typo for SWBD-DAMSL; the body text uses SWBD-DAMSL consistently elsewhere.","section":"§2.2"},{"comment":"The edge labels in the transition diagrams are difficult to parse; the text says green numbers are teaching and orange are tutoring, but the figure does not include a clear legend, and the notation '0.15 | -' is unexplained. A separate legend or column header would improve readability.","section":"Figure 8, Figure 9"},{"comment":"The captions contain 'T alkMove' and 'T alkMove' instead of 'TalkMove'; also 'Bigram Frequency Heatmap' is missing a space.","section":"Figure 10, Figure 11 captions"},{"comment":"The word 'explainations' in the contribution list should be 'explanations'.","section":"§1.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a descriptive study with a promising framework, but the central claim about non-talk-move utterances depends on unvalidated automatic labels. I do not see this as an immediate rejection because the authors explicitly acknowledge the limitation and the datasets and models are publicly available; however, the revision must either add a validation component or restructure the claims. I would also encourage the editor to consider whether the journal's audience expects inferential statistics for comparative claims between the two corpora; the current descriptive approach may be acceptable in a workshop paper but is a weakness at full-journal length."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nThis paper is a useful exploratory map, not a settled result. The genuinely new thing is the joint application of three annotation layers—human-annotated talk moves, SWBD-DAMSL dialogue acts, and SDRT discourse relations—to two math education corpora, teaching (TalkMoves) and tutoring (SAGA22). The comparative angle and the explicit focus on utterances that do not contain talk moves are worthwhile. The authors make a plausible case that T-NONE utterances are not just fillers, and the qualitative examples in Figures 15–17 do support that point.\n\nWhat the paper does well: the talk move labels are human-generated, giving solid ground for the frequency comparisons in Figure 2 and the transition diagrams. The limitations section is honest about the off-the-shelf models. The writing is clear and the framework is reproducible in principle.\n\nThe soft spots are real and load-bearing for the numbers. Every dialogue act and discourse relation label comes from models trained on Switchboard and Minecraft dialogue, and there is no validation on educational transcripts. That means the percentages in Table 3 and the DA distributions in Figure 7 are only as trustworthy as the cross-domain transfer assumption. The authors acknowledge this in Section 4.2, but the abstract and conclusion lean on those numbers without that caveat. The 10% transition threshold and the 5% frequency cutoff are arbitrary; no sensitivity analysis is offered. There are no confidence intervals or significance tests, so claims like 'significantly higher' or 'notable difference' are informal. The lack of released analysis code and an annotated validation sample makes it hard to replicate or probe.\n\nI do not think the circularity worry is the main issue—every utterance being tagged by the DA model makes the non-talk moves look functional by construction, but the paper's real point is about which functions dominate, which still depends on the labels. The qualitative examples save the central claim from being empty, but the quantified distributional claims do not stand alone.\n\nWho this is for: researchers in educational NLP and learning analytics, especially those building automated teacher feedback or AI tutoring agents. It deserves a serious referee—the idea is promising and the human-annotated talk move backbone is valuable. I would send it out with a request for either a validation sample with human agreement on the DA/DR labels, a filtered version of the analysis, or a reframing as an exploratory study with more cautious wording.","headline":"Promising multi-layer view of math teaching/tutoring dialogue, but the central claim about non-talk moves needs validation before the percentages are trusted.","tokens_in":22651,"tokens_out":2759,"would_cite":false,"duration_ms":26847,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that utterances lacking any coded talk move are not filler but actively guide, acknowledge, and structure mathematics classroom and tutoring discourse, and demonstrates this with a three-view analysis of teaching and…","keywords":["mathematics classroom discourse","talk moves","dialogue acts","discourse relations","SDRT","tutoring","automated feedback","educational NLP"],"falsifier":"Annotate a random sample of roughly 200 teacher and 200 student utterances from TalkMoves and SAGA22 with human dialogue-act and discourse-relation labels using the same schemas, then compare against the off-the-shelf model outputs; if agreement for the key acts (Statement-non-opinion, Acknowledge-(Backchannel), Action-directive) and relations (Continuation, Elaboration, Clarification_question) is low, or if human labels on math-specific references such as pointing to a drawing differ systematically from the models', the paper's distributional and sequential findings would not stand.","tokens_in":21539,"feed_emoji":"💬","tokens_out":8515,"duration_ms":78960,"temperature":0.7,"pith_summary":"This paper tries to establish that the roughly half of classroom and tutoring utterances that carry no theory-grounded talk move are not conversational filler, but active components of mathematics instruction that guide discussions, acknowledge student contributions, and maintain coherence. It supports this by laying three annotation views over the same transcripts—talk moves, 43 dialogue acts, and 16 discourse relations—and running unigram, sequential, and deep-dive analyses on two large transcript sets, one from classroom teaching and one from online tutoring. The payoff for a general reader is a concrete redesign target: automated feedback for educators and the behavior policies of AI tutoring agents should be built from all three views, not from talk-move detection alone. It also documents tutoring-specific gaps, such as tutors' higher rate of untagged utterances and lower rate of restating student ideas, which could become direct coaching feedback.","feed_headline":"Non-talk-move utterances are not filler—they structure math talk","feed_subtitle":"More than half of all utterances lack a talk move; dialogue acts and discourse relations reveal their function.","key_machinery":"The load-bearing mechanism is a triple annotation stack over one utterance stream: (1) domain-specific talk moves from Accountable Talk theory (seven teacher moves such as T-PRSACC and T-KPTG, five student moves such as S-MCLAIM and S-PROEVI); (2) a flattened multi-functional dialogue-act schema (SWBD-DAMSL, 43 tags) that assigns each utterance one mutually exclusive act such as Wh-Question, Action-directive, or Acknowledge-(Backchannel); and (3) Segmented Discourse Representation Theory (SDRT) with 16 discourse relations that connect an utterance to neighbors in a graph. The analysis pipeline runs top-down: unigram distributions of talk moves, sequential transition probabilities over talk-move bigrams with and without intervening non-talk utterances, and a multi-view deep dive that reads the dialogue acts and discourse relations attached to selected high-frequency bigrams. The off-the-shelf dialogue-act and discourse-relation parsers are what allow non-talk utterances to receive functional labels at all.","core_discovery":"On the paper's own terms, the central discovery is that utterances labeled T-NONE and S-NONE, which together make up more than half of all dialogue in both datasets, carry systematic communicative functions. When tagged with dialogue acts, they are predominantly Statement-non-opinion, Acknowledge-(Backchannel), Action-directive, and Continued-by-same-speaker; when linked by discourse relations, they enter Continuation, Elaboration, Clarification_question, and Acknowledgement relations with adjacent talk moves. The paper interprets this as evidence that these utterances guide discussions, acknowledge student input, and bridge or scaffold teacher and student moves, and it derives actionable contrasts between teaching and tutoring from the same evidence.","pith_inferences":["Editorial inference: if the transfer from general conversation models holds, the same three-view pipeline could be applied to new classroom corpora without retraining, making DA/DR labels a cheap complement to talk-move annotation.","Editorial inference: the paper's own reported numbers are consistent with the hypothesis that the talking-to-think mechanisms behind accountable talk operate through the untagged scaffolding utterances; a direct test would be to ablate T-NONE utterances from transcripts and see whether identifiable talk-move patterns or lesson quality degrade.","Editorial inference: the framework's value for AI agents is testable by building a tutor response generator conditioned on all three views and comparing student engagement against a talk-move-only baseline.","Editorial inference: because the flattened SWBD-DAMSL schema sacrifices DAMSL's multi-layer expressiveness, some multifunctionality may still be lost; re-annotating a sample with original multi-label DAMSL layers would show whether the flattened tags undercount utterances that both guide and acknowledge."],"forward_implications":["Automated feedback that reports only talk moves omits the guiding, acknowledging, and structuring functions that occupy most utterances; adding dialogue acts and discourse relations would give educators feedback on the full discourse.","The tutoring data's higher non-talk share and its lower restating rate give specific, measurable coaching targets for tutors: restate student ideas more often and reduce reliance on untagged directives.","For student talk moves like S-MCLAIM, the dominant discourse relations toward the teacher are Clarification_question, Continuation, and Elaboration; an AI agent designed to respond to student claims could be patterned on these observed relations.","High-probability transitions such as T-PRSREA to S-PROEVI (41% in tutoring) and T-PRSACC to S-MCLAIM identify the pedagogical routines that talk-move-based professional development should emphasize.","The finding that T-NONE utterances bridge same-category teacher talk moves implies that some 'single' moves are actually multi-utterance spans, which matters for how feedback should segment and aggregate moves."],"supporting_citations":[{"why":"supplies the TalkMoves classroom transcript dataset with human talk-move labels used throughout the analysis.","marker":"[64]"},{"why":"supplies the SAGA22 tutoring transcript dataset and human talk-move labels, and documents that over 50% of utterances are non-talk moves.","marker":"[18]"},{"why":"defines the flattened SWBD-DAMSL dialogue-act schema whose 42/43 tags are applied to every utterance.","marker":"[41]"},{"why":"provides the off-the-shelf speaker-aware dialogue-act tagger that assigns the DA labels.","marker":"[34]"},{"why":"defines SDRT and its discourse relations, the theory behind the third annotation view.","marker":"[5]"},{"why":"provides the off-the-shelf SDRT discourse parser used to label discourse relations.","marker":"[67]"},{"why":"is the dialogue corpus on which the discourse parser was trained, and so the source of the cross-domain transfer assumption.","marker":"[68]"},{"why":"establishes the prior finding that non-talk moves are prevalent and under-studied, motivating the paper's focus.","marker":"[38]"}],"fun_headline_variants":["Half of math dialogue lacks talk moves—but still guides learning","Non-talk moves: the hidden structure in math teaching and tutoring","Beyond talk moves: how dialogue acts reveal math discourse patterns","Talk moves miss half the action—dialogue acts fill the gap","Multi-perspective analysis uncovers roles of non-talk utterances"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that dialogue-act and discourse-relation labels produced by off-the-shelf models trained on general conversational data (telephone and game dialogue, not mathematics instruction) are accurate enough on TalkMoves and SAGA22 that every percentage, transition, and qualitative example in the analysis inherits their validity; the paper does not validate these labels on either dataset.","fun_headline_variants_meta":{"raw":{"variants":["Half of math dialogue lacks talk moves—but still guides learning","Non-talk moves: the hidden structure in math teaching and tutoring","Beyond talk moves: how dialogue acts reveal math discourse patterns","Talk moves miss half the action—dialogue acts fill the gap","Multi-perspective analysis uncovers roles of non-talk utterances"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000173,"raw_usage":{"total_tokens":1294,"prompt_tokens":977,"completion_tokens":317,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":593,"completion_tokens_details":{"reasoning_tokens":231}},"tokens_in":593,"tokens_out":317,"duration_ms":3252,"temperature":1.0,"reasoning_tokens":231,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:22:25.845022+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Annotate a random sample of roughly 200 teacher and 200 student utterances from TalkMoves and SAGA22 with human dialogue-act and discourse-relation labels using the same schemas, then compare against the off-the-shelf model outputs; if agreement for the key acts (Statement-non-opinion, Acknowledge-(Backchannel), Action-directive) and relations (Continuation, Elaboration, Clarification_question) is low, or if human labels on math-specific references such as pointing to a drawing differ systematically from the models', the paper's distributional and sequential findings would not stand.","supporting_citations":[{"cited_title":"Prasad, N","cited_arxiv_id":null,"evidence_quote":"supplies the TalkMoves classroom transcript dataset with human talk-move labels used throughout the analysis."},{"cited_title":"Boussioux, J","cited_arxiv_id":null,"evidence_quote":"supplies the SAGA22 tutoring transcript dataset and human talk-move labels, and documents that over 50% of utterances are non-talk moves."},{"cited_title":"Jurafsky","cited_arxiv_id":null,"evidence_quote":"defines the flattened SWBD-DAMSL dialogue-act schema whose 42/43 tags are applied to every utterance."},{"cited_title":"Demszky, J","cited_arxiv_id":null,"evidence_quote":"provides the off-the-shelf speaker-aware dialogue-act tagger that assigns the DA labels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines SDRT and its discourse relations, the theory behind the third annotation view."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the off-the-shelf SDRT discourse parser used to label discourse relations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"is the dialogue corpus on which the discourse parser was trained, and so the source of the cross-domain transfer assumption."}],"review_version":1}