Pith. sign in

REVIEW 4 major objections 5 minor 80 references

Towards Actionable Pedagogical Feedback: A Multi-Perspective Analysis of Mathematics Teaching and Tutoring Dialogue

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper claims that utterances lacking any coded talk move are not filler but actively guide, acknowledge, and structure mathematics classroom and tutoring discourse, and demonstrates this with a three-view analysis of teaching and…

desk verdict Promising multi-layer view of math teaching/tutoring dialogue, but the central claim about non-talk moves needs validation before the percentages are trusted. read the letter →

arxiv 2505.07161 v1 pith:WA2VQ4TI submitted 2025-05-12 cs.CL cs.AIcs.HCcs.LG

classification cs.CLcs.AIcs.HCcs.LG
keywords mathematicsclassroomdiscoursetalkmovesdialogueactsrelationsSDRTtutoringautomatedfeedbackeducationalNLP
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that the roughly half of classroom and tutoring utterances that carry no theory-grounded talk move are not conversational filler, but active components of mathematics instruction that guide discussions, acknowledge student contributions, and maintain coherence. It supports this by laying three annotation views over the same transcripts—talk moves, 43 dialogue acts, and 16 discourse relations—and running unigram, sequential, and deep-dive analyses on two large transcript sets, one from classroom teaching and one from online tutoring. The payoff for a general reader is a concrete redesign target: automated feedback for educators and the behavior policies of AI tutoring agents should be built from all three views, not from talk-move detection alone. It also documents tutoring-specific gaps, such as tutors' higher rate of untagged utterances and lower rate of restating student ideas, which could become direct coaching feedback.

What carries the argument

The load-bearing mechanism is a triple annotation stack over one utterance stream: (1) domain-specific talk moves from Accountable Talk theory (seven teacher moves such as T-PRSACC and T-KPTG, five student moves such as S-MCLAIM and S-PROEVI); (2) a flattened multi-functional dialogue-act schema (SWBD-DAMSL, 43 tags) that assigns each utterance one mutually exclusive act such as Wh-Question, Action-directive, or Acknowledge-(Backchannel); and (3) Segmented Discourse Representation Theory (SDRT) with 16 discourse relations that connect an utterance to neighbors in a graph. The analysis pipeline runs top-down: unigram distributions of talk moves, sequential transition probabilities over talk-move bigrams with and without intervening non-talk utterances, and a multi-view deep dive that reads the dialogue acts and discourse relations attached to selected high-frequency bigrams. The off-the-shelf dialogue-act and discourse-relation parsers are what allow non-talk utterances to receive functional labels at all.

What would settle it

Annotate a random sample of roughly 200 teacher and 200 student utterances from TalkMoves and SAGA22 with human dialogue-act and discourse-relation labels using the same schemas, then compare against the off-the-shelf model outputs; if agreement for the key acts (Statement-non-opinion, Acknowledge-(Backchannel), Action-directive) and relations (Continuation, Elaboration, Clarification_question) is low, or if human labels on math-specific references such as pointing to a drawing differ systematically from the models', the paper's distributional and sequential findings would not stand.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that utterances labeled T-NONE and S-NONE, which together make up more than half of all dialogue in both datasets, carry systematic communicative functions. When tagged with dialogue acts, they are predominantly Statement-non-opinion, Acknowledge-(Backchannel), Action-directive, and Continued-by-same-speaker; when linked by discourse relations, they enter Continuation, Elaboration, Clarification_question, and Acknowledgement relations with adjacent talk moves. The paper interprets this as evidence that these utterances guide discussions, acknowledge student input, and bridge or scaffold teacher and student moves, and it derives actionable contrasts between teaching and tutoring from the same evidence.

Load-bearing premise

The load-bearing premise is that dialogue-act and discourse-relation labels produced by off-the-shelf models trained on general conversational data (telephone and game dialogue, not mathematics instruction) are accurate enough on TalkMoves and SAGA22 that every percentage, transition, and qualitative example in the analysis inherits their validity; the paper does not validate these labels on either dataset.

Editorial extensions

If this is right

  • Automated feedback that reports only talk moves omits the guiding, acknowledging, and structuring functions that occupy most utterances; adding dialogue acts and discourse relations would give educators feedback on the full discourse.
  • The tutoring data's higher non-talk share and its lower restating rate give specific, measurable coaching targets for tutors: restate student ideas more often and reduce reliance on untagged directives.
  • For student talk moves like S-MCLAIM, the dominant discourse relations toward the teacher are Clarification_question, Continuation, and Elaboration; an AI agent designed to respond to student claims could be patterned on these observed relations.
  • High-probability transitions such as T-PRSREA to S-PROEVI (41% in tutoring) and T-PRSACC to S-MCLAIM identify the pedagogical routines that talk-move-based professional development should emphasize.
  • The finding that T-NONE utterances bridge same-category teacher talk moves implies that some 'single' moves are actually multi-utterance spans, which matters for how feedback should segment and aggregate moves.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: if the transfer from general conversation models holds, the same three-view pipeline could be applied to new classroom corpora without retraining, making DA/DR labels a cheap complement to talk-move annotation.
  • Editorial inference: the paper's own reported numbers are consistent with the hypothesis that the talking-to-think mechanisms behind accountable talk operate through the untagged scaffolding utterances; a direct test would be to ablate T-NONE utterances from transcripts and see whether identifiable talk-move patterns or lesson quality degrade.
  • Editorial inference: the framework's value for AI agents is testable by building a tutor response generator conditioned on all three views and comparing student engagement against a talk-move-only baseline.
  • Editorial inference: because the flattened SWBD-DAMSL schema sacrifices DAMSL's multi-layer expressiveness, some multifunctionality may still be lost; re-annotating a sample with original multi-label DAMSL layers would show whether the flattened tags undercount utterances that both guide and acknowledge.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a multi-perspective discourse analysis framework that combines domain-specific talk moves (human-annotated in the TalkMoves and SAGA22 datasets), automatic dialogue acts (SWBD-DAMSL-style, 43 tags), and automatic discourse relations (SDRT, 16 relations) to study mathematics teaching and tutoring dialogues. The authors perform unigram distributions, sequential talk-move transition analyses, and multi-view deep dives, and report that utterances without talk moves (T-NONE/S-NONE) are not mere fillers but contribute to guiding, acknowledging, and structuring classroom discourse. The stated goal is to provide actionable feedback for educators and design principles for AI agents in mathematics education.

Significance. The paper addresses a genuine and timely problem: most automated classroom-discourse feedback focuses only on theory-driven talk moves and ignores the majority of utterances that do not fit those moves. The integration of human-annotated talk moves with off-the-shelf dialogue act and discourse relation models over two large educational datasets is novel, and the open-sourced models and reproducible analysis pipeline are strengths. If the central claim about the role of non-talk-move utterances were supported by validated labels, the framework would be a useful step toward richer, action-oriented feedback and toward more capable AI tutoring agents. However, as presented, the main conclusion rests on automatic labels that have not been shown to be reliable on educational discourse, so the significance is conditional on that validation being supplied.

major comments (4)
  1. [§3.1.2, §3.3.2, Fig. 7, Table 3] The paper's central claim—that T-NONE utterances are not fillers but 'guide, acknowledge, and structure' discourse (abstract; §4.1)—is inferred directly from dialogue act and discourse relation distributions produced by models trained on Switchboard (the DA model, §2.2) and Minecraft dialogue (Llamipa, §2.2), neither of which was validated on TalkMoves or SAGA22. The limitation is acknowledged in §4.2, but the empirical results in Fig. 7 and Table 3 are nevertheless presented as evidence for the central claim. Without a human-annotated evaluation sample, an error analysis, or confidence-based filtering that demonstrates these cross-domain labels are accurate on mathematics classroom and tutoring transcripts, the quantitative foundation of the main conclusion is missing. The authors should either provide such validation or substantially soften the claims and mark them as hypotheses to be tested with in-domain labels.
  2. [§2.3.2, Eq. (1), Figs. 8–11] The sequential analysis depends on several ad-hoc thresholds: the 10% transition probability cutoff in Figs. 8 and 9, the 5% frequency exclusion for T-NONE counts in §2.3.2, and the top-3/top-7 DA selection in §2.3.1. No sensitivity analysis is provided, so it is unclear whether the reported patterns, such as the 'higher T-NONE interactions in tutoring' (§3.2.2), are robust to reasonable threshold changes. Additionally, Eq. (1) is under-specified: it appears to compute an expected number of intervening T-NONE utterances, but the treatment of zero T-NONE cases and the exclusion rule 'frequency below 5%' are not clearly defined. The authors should justify the thresholds or test several values.
  3. [§3.2.1, §3.1] All cross-domain comparisons (teaching vs. tutoring) are descriptive, yet the text makes claims such as 'significantly higher' for the T-PRSREA→S-PROEVI transition (41% vs. 28%, §3.2.1) and states differences in T-NONE prevalence (5.8% and 4.3%) as meaningful. The data have a nested structure (utterances within sessions) and the two datasets differ in session count, number of students per session, and session length, which may confound simple frequency comparisons. The paper should report confidence intervals, effect sizes, or a multilevel model; at minimum, the wording should be 'descriptively different' rather than 'significant'.
  4. [§2.3.1, Table 3] The unigram DA analysis explicitly excludes the Continued-by-same-speaker dialogue act (§2.3.1), yet Table 3 reports ContS. as the most frequent DA in several T-NONE transition contexts (e.g., 43.74% for T-RESTAT→T-NONE and 61.66% for S-MCLAIM→T-NONE). This inconsistency between the bottom-up DA analysis and the top-down transition analysis is confusing and should be resolved: either apply the same exclusion/inclusion criterion everywhere or explain why ContS. is meaningful in the transition analysis but excluded from unigram distributions.
minor comments (5)
  1. [Figure 2] The percentages in Figure 2 appear inconsistent with the text: the caption/legend lists T-None at 49.3% but the pie label shows 49.7%, and S-None is listed as 10.92% but labeled 11.0%; please reconcile.
  2. [§2.2] The phrase 'flattened multi-functional SWBD-MASL schema' in the abstract and §1.2 is likely a typo for SWBD-DAMSL; the body text uses SWBD-DAMSL consistently elsewhere.
  3. [Figure 8, Figure 9] The edge labels in the transition diagrams are difficult to parse; the text says green numbers are teaching and orange are tutoring, but the figure does not include a clear legend, and the notation '0.15 | -' is unexplained. A separate legend or column header would improve readability.
  4. [Figure 10, Figure 11 captions] The captions contain 'T alkMove' and 'T alkMove' instead of 'TalkMove'; also 'Bigram Frequency Heatmap' is missing a space.
  5. [§1.2] The word 'explainations' in the contribution list should be 'explanations'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: T-NONE findings are interpretive summaries of external model labels, not fitted parameters or self-citation chains.

full rationale

This is a descriptive, corpus-based empirical study rather than a derivational one, so most circularity patterns do not apply. The talk-move labels are human-annotated from TalkMoves and SAGA22, and the dialogue-act and discourse-relation labels are produced by off-the-shelf models (a SWBD-DAMSL dialogue-act classifier and the Llamipa SDRT parser) that were not trained or fine-tuned on these datasets. Therefore the percentages in Figures 3, 4, and 7 and Tables 2-3 are not fitted parameters renamed as predictions, and no equation reduces to its own input. The abstract and Section 4.1 conclude that T-NONE utterances 'guide, acknowledge, and structure' discourse; this is an ordinary-language summary of the DA tags (Action-directive, Acknowledgment-(Backchannel), Statement-non-opinion) and DR tags (Continuation, Elaboration, Comment) assigned to T-NONE, but it is not derivable by definition from the T-NONE label, which is defined only as the absence of the seven teacher talk moves. The main weakness is that the interpretative claim inherits the semantics (and potential errors) of the cross-domain models, which the authors explicitly acknowledge in Section 4.2 as a limitation; that is a validity or domain-transfer concern, not a circularity. Self-citations to prior papers for the datasets are references to independently human-annotated resources and are not used to justify the central interpretive claim through a self-citation chain. No step meets the evidentiary standard of reducing by construction to its own inputs, so no circularity is found.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

The paper introduces no new numerical fit to its own data and no new theoretical entities. Its load-bearing assumptions are the validity of human talk move labels and the transferability of two automatic annotation models. The free parameters are analysis thresholds, not scientific constants, but they shape every reported pattern.

free parameters (4)
  • transition probability threshold = 10%
    Chosen to filter transitions shown in Figures 8 and 9; arbitrary cutoff that determines which patterns are reported.
  • DA coverage threshold for talk moves = top 3 DAs, at least 50% coverage
    Section 2.3.1 selects top three DAs per talk move to ensure at least 50% of the distribution is captured; affects Figure 4.
  • DA coverage threshold for non-talk moves = top 7 DAs, at least 75% coverage
    Section 2.3.1 uses top seven DAs for non-talk move utterances to reach 75% coverage; affects Figure 7 and related analysis.
  • minimum T-NONE frequency for probability calculation = 5%
    Section 2.3.2 excludes T-NONE frequency counts below 5% when computing the probability of T-NONE occurring between talk move pairs; affects Figure 11.
assumptions (3)
  • domain assumption Human-annotated talk move labels in TalkMoves and SAGA22 are correct and reliable.
    The entire analysis treats these labels as ground truth, based on prior annotation work by Suresh et al. and Cao et al.; no inter-annotator agreement is reported in this paper.
  • ad hoc to paper Off-the-shelf SWBD-DAMSL dialogue act and SDRT discourse relation models transfer to mathematics educational dialogue.
    Section 2.2 uses models trained on Switchboard and Minecraft corpora; Section 4.2 admits they were not fine-tuned on educational data. All DA and DR results assume this transfer is valid enough for the reported patterns.
  • ad hoc to paper The selected analysis thresholds (10%, top-3/top-7, 5%) yield representative and meaningful patterns.
    These thresholds are arbitrary analysis choices, not derived from theory, and they determine the visible transitions and dialogue act summaries.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Actionable Pedagogical Feedback: A Multi-Perspective Analysis of Mathematics Teaching and Tutoring Dialogue." pith.science (2026). https://pith.science/paper/WA2VQ4TI

@misc{pith2026250507161,
  author       = {Pith},
  title        = {Pith review of: Towards Actionable Pedagogical Feedback: A Multi-Perspective Analysis of Mathematics Teaching and Tutoring Dialogue},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WA2VQ4TI}},
  note         = {Machine review of arXiv:2505.07161}
}
read the original abstract

Effective feedback is essential for refining instructional practices in mathematics education, and researchers often turn to advanced natural language processing (NLP) models to analyze classroom dialogues from multiple perspectives. However, utterance-level discourse analysis encounters two primary challenges: (1) multifunctionality, where a single utterance may serve multiple purposes that a single tag cannot capture, and (2) the exclusion of many utterances from domain-specific discourse move classifications, leading to their omission in feedback. To address these challenges, we proposed a multi-perspective discourse analysis that integrates domain-specific talk moves with dialogue act (using the flattened multi-functional SWBD-MASL schema with 43 tags) and discourse relation (applying Segmented Discourse Representation Theory with 16 relations). Our top-down analysis framework enables a comprehensive understanding of utterances that contain talk moves, as well as utterances that do not contain talk moves. This is applied to two mathematics education datasets: TalkMoves (teaching) and SAGA22 (tutoring). Through distributional unigram analysis, sequential talk move analysis, and multi-view deep dive, we discovered meaningful discourse patterns, and revealed the vital role of utterances without talk moves, demonstrating that these utterances, far from being mere fillers, serve crucial functions in guiding, acknowledging, and structuring classroom discourse. These insights underscore the importance of incorporating discourse relations and dialogue acts into AI-assisted education systems to enhance feedback and create more responsive learning environments. Our framework may prove helpful for providing human educator feedback, but also aiding in the development of AI agents that can effectively emulate the roles of both educators and students.

Figures

Figures reproduced from arXiv: 2505.07161 by the authors.

Figure 1
Figure 1. A running example for our multi-perspective analysis with talk moves, dialogue acts, and discourse relations. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 3
Figure 3. Distribution of DAs in the Teaching and Tutoring Datasets. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 2
Figure 2. Comparison of Teacher/Tutor and Student Talk Moves [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Comparison of Talk Moves with DAs. form of statements. For example, S-ASKMI being expressed as a Statement may reflect a student behavior of implicitly conveying their need for additional information by articulating their confu￾sion rather than making a direct request.…
Figure 5
Figure 5. Figure 5: Interesting DA Use Cases for Talk moves rence of this phenomenon in the tutoring dataset, as observed from [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Examples of teacher talk moves T-GSTUR, T-KPTG, T-PRSACC, and the student talk moves S-ASKMI, S-MCLAIM, that aligns with the dialogue act Yes-No-Questions. Our results suggest that DAs offer a more nuanced understanding of the behavioral patterns of students and teache…
Figure 9
Figure 9. Figure 9: Transition probabilities of one talk move being followed by [PITH_FULL_IMAGE:figures/full_fig_p007_9.png]
Figure 11
Figure 11. Figure 11: Comparison of Probability of T-NONE occurrence in between Talk Moves [PITH_FULL_IMAGE:figures/full_fig_p008_11.png]
Figure 14
Figure 14. Figure 14: Example of the S-MCLAIM → T-PRSACC talk move pair with the Clarification Question discourse relation and the S￾MCLAIM → T-REVOIC talk move pair with the Continuation, Elab￾oration, and Acknowledgement discourse relations. Comment. The most frequent DAs associated with…
Figure 12
Figure 12. Figure 12: Examples of the S-RELTO- S-RELTO with the Continu￾ation, Contrast, and Acknowledgement discourse relations. S: Yes, when you divide you make the number smaller. S: You move the decimal in front of the number you started with. Elaboration Micha: Well, if every time the…
Figure 13
Figure 13. Figure 13: Examples of the S-PROEVI- S-PROEVI talk move pair with the Elaboration, Continuation, Correction, and Contrast dis￾course relations. 3.3.2 Importance of Utterances without Talk Moves We conducted an in-depth analysis of discursive interactions not involving talk moves…
Figure 16
Figure 16. Figure 16: Examples of the same-category talk moves separated by [PITH_FULL_IMAGE:figures/full_fig_p010_16.png]
Figure 17
Figure 17. Figure 17: Examples of different-category talk moves separated by [PITH_FULL_IMAGE:figures/full_fig_p010_17.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

80 extracted references · 74 canonical work pages

  1. [1]

    INTRODUCTION Research increasingly supports dialogic teaching (learning) - an approach that encourages student-driven academic discourse — as a means of enhancing motivation and learning outcomes [10, 36, 60, 71]. A critical element of dialogic instruction is accountable talk, which structures classroom discourse around three key dimensions: learning comm...

  2. [2]

    transitions

    METHODS 2.1 Dataset We utilized two primary datasets from mathematics K-12 educa- tional contexts, one teaching dataset TalkMoves and one tutoring dataset SAGA22, summarized in Table 1. Each session’s transcript from the two datasets has been meticulously annotated by human experts, labeling instances of 7 distinct teacher talk moves and 5 student talk mo...

  3. [3]

    Bye" and

    RESULTS 3.1 Results from Unigram Analysis By examining the unigram distribution of teacher and student talk moves in both the teaching and tutoring datasets, Figure 2 shows 4 49.7% 11.0% 9.8% 9.7% 7.8% T-None (49.3%) S-None (10.92%) T-PrsAcc (9.71%) T-KpTg (9.65%) S-Mclaim (7.75%) S-ProEvi (3.38%) S-RelTo (2.8%) T-Revoic (1.67%) T-GStuR (1.21%) T-Restat (...

  4. [4]

    under the hood

    DISCUSSION We analyzed pedagogical behaviors in two mathematics education datasets, TalkMoves(teaching) and SAGA22 (tutoring). Using man- ually annotated talk moves and two state-of-the-art models for di- alogue act (DAs) and discourse relation (DRs) prediction, we con- ducted a top-down analysis on discursive behaviors looking at (1) unigram patterns, (2...

  5. [5]

    CONCLUSION Integrating talk moves, dialogue acts, and discourse relations, our multi-perspective study reveals key insights into the nature of teach- ing and tutoring discourse across two mathematic education datasets: TalkMoves and SAGA22. Our analysis reveals that while teaching and tutoring datasets share overarching distributional similarities in thei...

  6. [6]

    This research was supported by the National Science Foundation grant #2222647 and the NSF National AI In- stitute for Student-AI Teaming (iSAT) under grant DRL #2019805

    ACKNOWLEDGMENTS The authors would like to thank the anonymous reviewers for their valuable feedback. This research was supported by the National Science Foundation grant #2222647 and the NSF National AI In- stitute for Student-AI Teaming (iSAT) under grant DRL #2019805. All opinions are those of the authors and do not reflect those of the funding agencies

  7. [7]

    Afantenos, N

    S. Afantenos, N. Asher, F. Benamara, M. Bras, C. Fabre, M. Ho-dac, A. L. Draoulec, P. Muller, M.-P. Péry-Woodley, L. Prévot, J. Rebeyrolles, L. Tanguy, M. Vergez-Couret, and L. Vieu. An empirical resource for discovering cognitive principles of discourse organisation: the ANNODIS corpus. In N. Calzolari, K. Choukri, T. Declerck, M. U. Do˘gan, B. Maegaard,...

  8. [8]

    Allen and M

    J. Allen and M. Core. Draft of damsl: Dialog act markup in several layers, 1997

Show all 80 references
  1. [9]

    J. Allwood. An activity based approach to pragmatics. 1995

  2. [10]

    Asher, J

    N. Asher, J. Hunter, M. Morey, B. Farah, and S. Afantenos. Discourse structure and dialogue acts in multiparty dialogue: the STAC corpus. InProceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 2721–2727, Portorož, Slovenia, Ma...

  3. [11]

    Asher and A

    N. Asher and A. Lascarides. Logics of conversation. Cambridge University Press, 2003

  4. [12]

    J. L. Austin. How to do things with words. Oxford university press, 1975

  5. [13]

    Benamara and M

    F. Benamara and M. Taboada. Mapping different rhetorical relation annotations: A proposal. In M. Palmer, G. Boleda, and P. Rosso, editors, Proceedings of the Fourth Joint Conference on Lexical and Computational Semantics, pages 147–152, Denver, Colorado, June 2015. Association...

  6. [14]

    Bennis, J

    Z. Bennis, J. Hunter, and N. Asher. A simple but effective model for attachment in discourse parsing with multi-task learning for relation labeling. In A. Vlachos and I. Augenstein, editors, Proceedings of the 17th Conference of the European Chapter of the Association for Comp...

  7. [15]

    M. Z. bin Mohamed, R. Hidayat, N. N. binti Suhaizi, M. K. H. bin Mahmud, S. N. binti Baharuddin, et al. Artificial intelligence in mathematics education: A systematic literature review. International Electronic Journal of Mathematics Education, 17(3):em0694, 2022

  8. [16]

    Böheim, K

    R. Böheim, K. Schnitzler, A. Gröschner, M. Weil, M. Knogler, A.-K. Schindler, M. Alles, and T. Seidel. How changes in teachers’ dialogic discourse practice relate to changes in students’ activation, motivation and cognitive engagement. Learning, Culture and Social Interaction,...

  9. [17]

    B. M. Booth, J. Jacobs, J. B. Bush, B. Milne, T. Fischaber, and S. K. DMello. Human-tutor coaching technology (htct): Automated discourse analytics in a coached tutoring model. In Proceedings of the 14th Learning Analytics and Knowledge Conference, pages 725–735, 2024

  10. [18]

    Boussioux, J

    L. Boussioux, J. N. Lane, M. Zhang, V . Jacimovic, and K. R. Lakhani. The crowdless future? generative ai and creative problem-solving. Organization Science, 35(5):1589–1607, 2024

  11. [19]

    Breideband, J

    T. Breideband, J. Bush, C. Chandler, M. Chang, R. Dickler, P. Foltz, A. Ganesh, R. Lieber, W. R. Penuel, J. G. Reitman, et al. The community builder (cobi): Helping students to develop better small group collaborative learning skills. In Companion Publication of the 2023 Confe...

  12. [20]

    J. Cai, B. D. King, M. Perkoff, S. Dudy, J. Cao, M. Grace, N. Wojarnik, G. Ananya, J. Martin, M. Palmer, M. Walker, and J. Flanigan. Dependency dialogue acts — annotation scheme and case study. The 13th International Workshop on Spoken Dialogue Systems Technology, 2022

  13. [21]

    Calcagni, F

    E. Calcagni, F. Ahmed, A. L. Trigo-Clapés, R. Kershner, and S. Hennessy. Developing dialogic classroom practices through supporting professional agency: Teachers’ experiences of using the t-seda practitioner-led inquiry approach. Teaching and Teacher Education, 126:104067, 2023

  14. [22]

    J. Cao, R. Dickler, M. Grace, J. B. Bush, A. Roncone, L. M. Hirshfield, M. A. Walker, and M. S. Palmer. Designing an ai partner for jigsaw classrooms. Los Angeles, California., 2023

  15. [23]

    J. Cao, A. Ganesh, J. Cai, R. Southwell, E. M. Perkoff, M. Regan, K. Kann, J. H. Martin, M. Palmer, and S. D’Mello. A comparative analysis of automatic speech recognition errors in small group classroom discourse. In Proceedings of the 31st ACM Conference on User Modeling, Ada...

  16. [24]

    J. Cao, A. Suresh, J. Jacobs, C. Clevenger, A. Howard, C. Brown, B. Milne, T. Fischaber, T. Sumner, and J. H. Martin. Enhancing talk moves analysis in mathematics tutoring through classroom teaching discourse. In O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugeni...

  17. [25]

    Chi and A

    T.-C. Chi and A. Rudnicky. Structured dialogue discourse parsing. In O. Lemon, D. Hakkani-Tur, J. J. Li, A. Ashrafzadeh, D. H. Garcia, M. Alikhani, D. Vandyke, and O. Dušek, editors, Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue...

  18. [26]

    P. R. Cohen, J. L. Morgan, and M. E. Pollack. Intentions in communication. MIT press, 1990

  19. [27]

    M. G. Core and J. Allen. Coding dialogs with the damsl annotation scheme. In AAAI fall symposium on communicative action in humans and machines, volume 56, pages 28–35. Boston, MA, 1997

  20. [28]

    Correnti, L

    R. Correnti, L. C. Matsumura, M. Walsh, D. Zook-Howell, D. D. Bickel, and B. Yu. Effects of online content-focused coaching on discussion quality and reading achievement: Building theory for how coaching develops teachers’ adaptive expertise. Reading Research Quarterly, 56(3):...

  21. [29]

    Correnti, M

    R. Correnti, M. K. Stein, M. S. Smith, J. Scherrer, M. McKeown, J. Greeno, and K. Ashley. Improving teaching at scale: Design for the scientific measurement and learning of discourse practice. Socializing Intelligence Through Academic Talk and Dialogue. AERA, 284, 2015

  22. [30]

    Datta, J

    D. Datta, J. P. Bywater, M. Phillips, S. Lilly, J. L. Chiu, G. S. Watson, and D. E. Brown. Classifying mathematics teacher questions to support mathematical discourse. In International Conference on Artificial Intelligence in Education, pages 372–377. Springer, 2023

  23. [31]

    Demszky and H

    D. Demszky and H. Hill. The ncte transcripts: A dataset of elementary math classroom transcripts. arXiv preprint arXiv:2211.11772, 2022

  24. [32]

    Demszky and J

    D. Demszky and J. Liu. M-powering teachers: Natural language processing powered feedback improves 1: 1 instruction and student outcomes. In Proceedings of the Tenth ACM Conference on Learning@ Scale, pages 59–69, 2023

  25. [33]

    Demszky, J

    D. Demszky, J. Liu, H. C. Hill, D. Jurafsky, and C. Piech. Can automated feedback improve teachers’ uptake of student ideas? evidence from a randomized controlled trial in a large-scale online course. Educational Evaluation and Policy Analysis, 46(3):483–505, 2024

  26. [34]

    Demszky, J

    D. Demszky, J. Liu, H. C. Hill, S. Sanghi, and A. Chung. 12 Automated feedback improves teachers’ questioning quality in brick-and-mortar classrooms: Opportunities for further enhancement. Computers & Education, 227:105183, 2025

  27. [35]

    P. J. Donnelly, N. Blanchard, B. Samei, A. M. Olney, X. Sun, B. Ward, S. Kelly, M. Nystran, and S. K. D’Mello. Automatic teacher modeling from live classroom audio. In Proceedings of the 2016 conference on user modeling adaptation and personalization, pages 45–53, 2016

  28. [36]

    R. K. Franklin, J. O’Neill Mitchell, K. S. Walters, B. Livingston, M. B. Lineberger, C. Putman, R. Yarborough, and L. Karges-Bone. Using swivl robotic technology in teacher education preparation: A pilot study. TechTrends, 62:184–189, 2018

  29. [37]

    Y . Fu. Towards unification of discourse annotation frameworks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop, pages 132–142, Dublin, Ireland, May 2022. Association for Computational Linguistics

  30. [38]

    J. J. Godfrey, E. C. Holliman, and J. McDaniel. Switchboard: Telephone speech corpus for research and development. In Acoustics, speech, and signal processing, ieee international conference on, volume 1, pages 517–520. IEEE Computer Society, 1992

  31. [39]

    M. Hancher. The classification of cooperative illocutionary acts1. Language in society, 8(1):1–14, 1979

  32. [40]

    Z. He, L. Tavabi, K. Lerman, and M. Soleymani. Speaker turn modeling for dialogue act classification. In M.-F. Moens, X. Huang, L. Specia, and S. W.-t. Yih, editors, Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2150–2157, Punta Cana, Dominican R...

  33. [41]

    Jurafsky

    D. Jurafsky. Switchboard swbd-damsl shallow-discourse-function annotation coders manual. Institute of Cognitive Science Technical Report, 1997

  34. [42]

    J. R. Hobbs. Literature and cognition. Number 21. Center for the Study of Language (CSLI), 1990

  35. [43]

    C. Howe, S. Hennessy, N. Mercer, M. Vrikki, and L. Wheatley. Teacher–student dialogue during classroom teaching: Does it really impact on student outcomes? Journal of the learning sciences, 28(4-5):462–512, 2019

  36. [44]

    A. Y . Huang, O. H. Lu, and S. J. Yang. Effects of artificial intelligence–enabled personalized recommendations on learners’ learning engagement, motivation, and outcomes in a flipped classroom. Computers & Education, 194:104684, 2023

  37. [45]

    Jacobs, K

    J. Jacobs, K. Scornavacco, C. Harty, A. Suresh, V . Lai, and T. Sumner. Promoting rich discussions in mathematics classrooms: Using personalized, automated feedback to support reflection and instructional change. Teaching and Teacher Education, 112:103631, 2022

  38. [46]

    Jensen, M

    E. Jensen, M. Dale, P. J. Donnelly, C. Stone, S. Kelly, A. Godley, and S. K. D’Mello. Toward automated feedback on teacher discourse to enhance teacher learning. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, pages 1–13, 2020

  39. [47]

    Jensen, M

    E. Jensen, M. Dale, P. J. Donnelly, C. Stone, S. Kelly, A. Godley, and S. K. D’Mello. Toward Automated Feedback on Teacher Discourse to Enhance Teacher Learning. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, pages 1–13, 2020

  40. [48]

    Y . Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V . Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019

  41. [49]

    H. Kamp, J. v. Genabith, and U. Reyle. Discourse representation theory. In Handbook of philosophical logic, pages 125–394. Springer, 2011

  42. [50]

    J. Kim, H. Lee, and Y . H. Cho. Learning design to support student-ai collaboration: Perspectives of leading teachers for ai in education. Education and Information Technologies, 27(5):6069–6104, 2022

  43. [51]

    A. Latham. Conversational Intelligent Tutoring Systems: The State of the Art. In A. E. Smith, editor, Women in Computational Intelligence: Key Advances and Perspectives on Emerging Topics, Women in Engineering and Science, pages 77–101. Springer International Publishing, Cham, 2022

  44. [52]

    Lefstein and J

    A. Lefstein and J. Snell. Better than best practice: Developing teaching and learning through dialogue. Routledge, 2013

  45. [53]

    C. Li, C. Braud, M. Amblard, and G. Carenini. Discourse relation prediction and discourse parsing in dialogues with minimal supervision. In M. Strube, C. Braud, C. Hardmeier, J. J. Li, S. Loaiciga, A. Zeldes, and C. Li, editors, Proceedings of the 5th Workshop on Computational...

  46. [54]

    J. Li, M. Liu, M.-Y . Kan, Z. Zheng, Z. Wang, W. Lei, T. Liu, and B. Qin. Molweni: A challenge multiparty dialogues-based machine reading comprehension dataset with discourse structure. In D. Scott, N. Bel, and C. Zong, editors, Proceedings of the 28th International Conference...

  47. [55]

    O’Connor and S

    C. O’Connor and S. Michaels. Supporting teachers in taking up productive talk moves: The long road to professional learning at scale. International Journal of Educational Research, 97:166–175, 2019

  48. [56]

    W. C. Mann and S. A. Thompson. Rhetorical structure theory: Toward a functional theory of text organization. Text-interdisciplinary Journal for the Study of Discourse, 8(3):243–281, 1988

  49. [57]

    at-scale

    L. C. Matsumura, H. E. Garnier, S. C. Slater, and M. D. Boston. Toward measuring instructional interactions “at-scale”. Educational Assessment, 13(4):267–300, 2008

  50. [58]

    Mercer, R

    N. Mercer, R. Wegerif, and L. C. Major. The Routledge international handbook of research on dialogic education. Routledge Abingdon, 2019

  51. [59]

    Michaels, M

    S. Michaels, M. W. Hall, and L. B. Resnick. Accountable talk sourcebook: For classroom conversation that works. University of Pittsburgh Pittsburgh, PA, 2013

  52. [60]

    Michaels, M

    S. Michaels, M. C. O’Connor, M. W. Hall, and L. B. Resnick. Accountable talk® sourcebook. Pittsburg, PA: Institute for Learning University of Pittsburgh. Murphy, PK, Wilkinson, IAG, Soter, AO, Hennessey, MN, & Alexander, JF, 2010

  53. [61]

    Moreau-Pernet, Y

    B. Moreau-Pernet, Y . Tian, S. Sawaya, P. Foltz, J. Cao, B. Milne, and T. Christie. Classifying tutor discursive moves at scale in mathematics classrooms with large language models. In Proceedings of the Eleventh ACM Conference on Learning @ Scale, L@S ’24, page 361–365, New Y...

  54. [62]

    Stolcke, K

    A. Stolcke, K. Ries, N. Coccaro, E. Shriberg, R. Bates, D. Jurafsky, P. Taylor, R. Martin, C. V . Ess-Dykema, and M. Meteer. Dialogue act modeling for automatic tagging and recognition of conversational speech. Computational linguistics, 26(3):339–373, 2000

  55. [63]

    E. M. Perkoff, A. Bhattacharyya, J. Cai, and J. Cao. Comparing neural question generation architectures for reading comprehension. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational 13 Applications (BEA 2023), pages 556–566, 2023

  56. [64]

    Prasad, N

    R. Prasad, N. Dinesh, A. Lee, E. Miltsakaki, L. Robaldo, A. Joshi, and B. Webber. The penn discourse treebank 2.0. In Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC’08), 2008

  57. [65]

    J. L. Ramos, A. A. Cattaneo, F. P. de Jong, and R. G. Espadeiro. Pedagogical models for the facilitation of teacher professional development via video-supported collaborative learning. a review of the state of the art. Journal of Research on Technology in Education, 54(5):695–...

  58. [66]

    Reese, J

    B. Reese, J. Hunter, N. Asher, P. Denis, and J. Baldridge. Reference manual for the analysis and annotation of rhetorical structure. PhD thesis, University of Texas at Austin, 2007

  59. [67]

    L. B. Resnick, C. S. C. Asterhan, S. N. Clarke, and F. Schantz. Next Generation Research in Dialogic Learning, chapter 13, pages 323–338. John Wiley & Sons, Ltd, 2018

  60. [68]

    J. R. Searle and J. R. Searle. Speech acts: An essay in the philosophy of language, volume 626. Cambridge university press, 1969

  61. [69]

    N. Tran, B. Pierce, D. Litman, R. Correnti, L. C. Matsumura, et al. Multi-dimensional performance analysis of large language models for classroom discussion assessment. Journal of Educational Data Mining, 16(2):304–335, 2024

  62. [70]

    Suresh, J

    A. Suresh, J. Jacobs, C. Clevenger, V . Lai, C. Tan, J. H. Martin, and T. Sumner. Using ai to promote equitable classroom discussions: The talkmoves application. In Artificial Intelligence in Education: 22nd International Conference, AIED 2021, Utrecht, The Netherlands, June 1...

  63. [71]

    Suresh, J

    A. Suresh, J. Jacobs, C. Harty, M. Perkoff, J. H. Martin, and T. Sumner. The talkmoves dataset: K-12 mathematics lesson transcripts annotated for teacher and student discursive moves. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 4654–4662, 2022

  64. [72]

    Suresh, J

    A. Suresh, J. Jacobs, M. Perkoff, J. H. Martin, and T. Sumner. Fine-tuning transformers with additional context to classify discursive moves in mathematics classrooms. In Proceedings of the 17th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2022)...

  65. [73]

    Tack and C

    A. Tack and C. Piech. The ai teacher test: Measuring the pedagogical ability of blender and gpt-3 in educational dialogues. arXiv preprint arXiv:2205.07540, 2022

  66. [74]

    Thompson, A

    K. Thompson, A. Chaturvedi, J. Hunter, and N. Asher. Llamipa: An incremental discourse parser. In Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, editors,Findings of the Association for Computational Linguistics: EMNLP 2024, pages 6418–6430, Miami, Florida, USA, Nov. 2024. Associa...

  67. [75]

    Thompson, J

    K. Thompson, J. Hunter, and N. Asher. Discourse structure for the Minecraft corpus. In N. Calzolari, M.-Y . Kan, V . Hoste, A. Lenci, S. Sakti, and N. Xue, editors, Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Eval...

  68. [77]

    D. Wang, D. Shan, Y . Zheng, K. Guo, G. Chen, and Y . Lu. Can chatgpt detect student talk moves in classroom discourse? a preliminary comparison with bert. 2023

  69. [78]

    N. M. Webb, M. L. Franke, M. Ing, N. C. Johnson, and J. Zimmerman. The details matter in mathematics classroom dialogue. In L. M. Neil Mercer, Rupert Wegerif, editor,The Routledge international handbook of research on dialogic education, pages 530–546. Routledge, 2019

  70. [79]

    Wittgenstein

    L. Wittgenstein. Philosophical investigations. John Wiley & Sons, 2010

  71. [80]

    Z. Wu, D. Ji, K. Yu, X. Zeng, D. Wu, and M. Shidujaman. Ai creativity and the human-ai co-creation model. In Human-Computer Interaction. Theory, Methods and Tools: Thematic Area, HCI 2021, Held as Part of the 23rd HCI International Conference, HCII 2021, Virtual Event, July 24...

  72. [2021]

    Association for Computational Linguistics

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.