REVIEW 4 major objections 5 minor 29 references
Interaction Analysis by Humans and AI: A Comparative Perspective
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper argues that an LLM-based transcription, translation, and emotion-annotation pipeline makes Finnish child-interaction data analysable by non-Finnish researchers, and that its first results show mixed reality eliciting more…
desk verdict Useful LLM-annotation pipeline study undermined by an abstract that contradicts its own data; the MR-vs-Zoom claim is unsupported, but the pipeline evaluation deserves a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the five-stage annotation pipeline: cloud speech-to-text transcribes the Finnish audio; GPT-3.5 corrects grammar; GPT-4 performs context-aware refinement and translation; DeepL adds translation support; and then human annotators plus the two models label speaker turns and emotions using an eight-emotion wheel. Agreement is measured with the kappa coefficient, with set-overlap similarity as a secondary metric, and the benchmark annotator is 'Finn-top,' the Finnish-speaking annotator with the highest agreement against the other Finnish annotators. What this pipeline does is convert a language-barrier problem and a time problem into a single automated workflow whose outputs can be compared across two very different recording setups.
What would settle it
Have Finnish-speaking human annotators label the original recordings directly, without machine transcription or translation, and compare the MR-versus-Zoom emotion proportions with the LLM-pipeline proportions; if the human-only proportions show no MR advantage, the headline finding is an artifact of the pipeline. A simpler check is to compute transcription error separately for the MR and Zoom audio and see whether the condition with more errors is also the one with more emotion labels.
Extended reading notes
Core claim
On its own terms, the paper's discovery has two parts. Methodologically, it shows that cloud speech-to-text (overall character error rate 0.24, reduced to about 0.18 by LLM correction) followed by GPT-3.5, GPT-4, and DeepL translation lets an English-speaking annotator label Finnish child-interaction data, with sentiment agreement comparable to the average Finnish-speaking human annotator in two-person dialogues and a 32-hour human workload compressed to minutes. Substantively, using an eight-emotion annotation scheme, the emotion labels show that the mixed-reality sessions had a larger emotion distribution per session than the Zoom sessions, which the authors interpret as greater emotional engagement, even though the Zoom group had the highest share of positive labels (84.21% versus 76.40% for MR). The authors explicitly frame these as initial findings and note in the conclusion that MR did not surpass Zoom in promoting positive engagement, partly because technical bugs in the MR system may have disrupted or frustrated children.
Load-bearing premise
The comparison assumes that the AI-generated emotion labels reflect the children's actual feelings equally well in both conditions, even though the two conditions were recorded with different microphones and the transcripts still contained many transcription errors.
Editorial extensions
If this is right
- Non-Finnish researchers can analyse Finnish child-interaction data directly, with two-person sentiment annotation agreement close to that of Finnish-speaking human annotators.
- The transcription, correction, and LLM annotation steps work best in two-person dialogues, so the workflow suits paired collaborative tasks better than overlapping multi-child group discussions.
- If the emotion-label distributions are taken at face value, mixed reality can support emotionally expressive distributed collaboration among children, making it a candidate medium for remote classroom activities.
- Because the Zoom condition showed the highest share of positive labels while MR showed a larger overall emotional output, engagement comparisons depend on whether one counts proportions or volumes; both metrics are reported.
Reading between the lines
- Editorial inference: the same pipeline could be carried to other low-resource or child-speech-heavy languages, provided a small human-transcribed sample is kept to measure transcription error and calibrate the models.
- Editorial inference: a decisive follow-up would re-run the emotion analysis on fully human-corrected transcripts; if the MR versus Zoom difference in emotion-label volume persists, it is a property of the medium rather than of transcription noise.
- Editorial inference: the authors' own caveat that MR did not surpass Zoom on positive engagement, despite richer emotional expression, suggests that current MR implementation bugs may be masking the medium's genuine effect, which live holoportation rather than avatar-based MR could test.
- Editorial inference: the reported LLM tendency to over-label dominant speakers implies that future automated speaker analysis should combine diarization with lexical cues before emotion proportions are trusted in group settings.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper compares children's communication during a gesture-based guessing game in two conditions: Mixed Reality (MR, via HoloLens plus 3D camera/TV) and 2D video conferencing (Zoom). Audio-video data were transcribed with Google Cloud, corrected and translated using GPT-3.5, GPT-4, and DeepL, and then annotated for speaker identity and emotion (Plutchik's model) by human annotators and the same LLMs. The authors report time/cost savings from the LLM pipeline and evaluate inter-annotator agreement. The abstract's central claim is that MR fosters richer interaction, evidenced by higher emotional expression and heightened engagement, while also noting limitations in annotation accuracy.
Significance. If the central claim were supported, the finding would be valuable for designers of distributed collaborative learning environments for children, suggesting that MR may enhance emotional engagement. The paper also contributes a practical demonstration that LLM-based transcription, translation, and annotation can reduce human annotation effort and enable non-native speakers to analyze Finnish-language interaction data. These efficiency results are useful, though their validity depends on the reliability of the pipeline, which the paper itself acknowledges is limited (CER 0.18 after correction, kappa values ranging widely). The significance of the paper is substantially weakened by the fact that its primary MR-versus-Zoom claim is contradicted by the very data it reports.
major comments (4)
- [Abstract and Section 5.2 (Figure 2) and Section 6] The abstract states that 'MR fosters richer interaction, evidenced by higher emotional expression during annotation, and heightened engagement.' This is directly contradicted by the paper's own data: Section 5.2 reports that the Zoom group has 84.21% positive emotions versus 76.40% for MR, and Section 6 concludes that MR 'did not surpass Zoom in promoting positive engagement.' The only alternative evidence gestured at is 'the larger area under the MR curve,' but no raw counts, confidence intervals, or significance tests are provided for that claim. Unless 'richer interaction' is explicitly redefined (e.g., as total number of emotion labels per session), the primary claim is unsupported by the reported results.
- [Section 5.1 (Table 5)] The text claims that 'GPT-4 outperformed both GPT-3.5 and the En-speaker in the experimental condition (κ = 0.4353 vs. 0.5995/0.752).' However, the numbers show the opposite: GPT-4's kappa of 0.4353 is lower than GPT-3.5's 0.5995 and En-speaker's 0.752. Table 5 in Appendix A.3 confirms this ordering. This is a load-bearing error because the section's conclusion about GPT-4's superiority in complex scenes depends on this reversed comparison.
- [Sections 4 and 5.2 (MR vs. Zoom comparison)] The comparison of emotional engagement between MR and Zoom relies entirely on emotion annotations produced by the same LLM pipeline being evaluated. The transcription error rate is substantial (CER 0.24 before correction, 0.18 after), and audio capture differs across conditions: MR uses dedicated microphones while Zoom audio is used for the control group. These differences could systematically distort emotion annotations in a condition-specific way (e.g., differing audio quality, speaker overlap, or translator behavior). The paper provides no analysis demonstrating that measurement error is non-differential across conditions, so the platform comparison is confounded with pipeline accuracy. This threat is acknowledged only as a general limitation, not addressed for the central comparison.
- [Section 5.2 (emotion percentages and normalization)] The percent-positive comparison is based on 'scaled emotion counts normalized by the session count' (4 interviews, 3 Zoom, 5 MR). This normalization presumes that each session contributes equally and that the number of emotion labels per session is not itself a meaningful outcome. Yet the claim of 'greater emotional engagement' appears to rest on the larger total area under the MR curve, which is exactly the unnormalized count. No statistical test or confidence interval accompanies either the percentage comparison or the area-under-curve assertion. With only 5 MR pairs and 3 Zoom pairs, the reported differences (84.21% vs. 76.40%) may not be reliable, and the paper should present at least a test of the difference or a clear statement that the difference is descriptive only.
minor comments (5)
- [Section 5.1] The kappa values reported in the running text (e.g., 'Finn-top achieved the highest agreement in both interview (κ = 0.5348) and experimental (κ = 0.928)') are presented in an order that is easy to misread; a consistent table-first presentation would help.
- [Figure 2 caption] The caption says 'scaled emotion counts normalized by the session count,' but the text in Section 5.2 refers to 'the larger area under the MR curve.' The figure does not show raw counts, so the area interpretation is not directly visible. Please clarify whether the plotted values are per-session averages or totals.
- [Section 4 (Initial Data)] The data description says '5 pairs, with 3 pairs in the control group' and '17 files and 250 minutes of Finnish-language data.' It would be helpful to state explicitly how many MR pairs (presumably 2) and how the 3 vs. 2 imbalance is handled in the analyses.
- [Section 5.2 (Sentiment agreement)] The sentence 'GPT-3.5 outperforms GPT-4 overall (GPT-4 interview: κ = 0.5943)' is confusing because 'overall' is not defined; the immediately preceding kappas are for Zoom, MR, and interviews separately. Consider stating a combined or averaged score if that is intended.
- [Abstract and Section 1] The phrase 'higher emotional expression during annotation' is ambiguous: it could mean the annotators expressed emotions, or the children's emotional expressions as captured in annotations. Rephrase to avoid this ambiguity.
Circularity Check
No circular derivation found: the MR-vs-Zoom conclusion is empirically contradicted by the paper's own figures, not constructed from its inputs.
full rationale
The paper contains no derivation step in which a predicted quantity is equivalent to a fitted input by construction. The central comparison between MR and Zoom is based on measured emotion annotations (Section 5.2), not on a parameter fitted to the outcome. The abstract's claim that 'MR fosters richer interaction' is indeed unsupported, and internally contradicted, by the paper's own reported positive-emotion shares ('The Interview Group contains 71.45% positive emotions, the MR group 76.40% positive, and the Zoom group 84.21% positive') and by Section 6's admission that MR 'did not surpass Zoom in promoting positive engagement,' but that is a correctness and consistency problem, not circularity. The one visible self-citation, reference [10] in Section 1, is used only to motivate the cost of manual transcription and is not load-bearing for any result. The LLM pipeline does create a methodological confound: the English-speaking annotator works from LLM-translated text, so the claimed cross-language validation is partly mediated by the very models under evaluation. However, no reported number reduces to an earlier fitted input, and the acknowledged limitations (CER 0.18–0.24, sentiment kappas roughly 0.56–0.93) are stated as limitations rather than presented as independent validation. A reporting error in Section 5.1 (GPT-4's kappa 0.4353 described as outperforming GPT-3.5's 0.5995) is also flagged as an internal inconsistency, not as a circular step. Under the requirement that circularity be exhibited by a specific reduction, this paper's central problems are evidentiary and logical, not circular.
Assumptions & free parameters
assumptions (5)
- domain assumption Plutchik's wheel of emotions (8 basic emotions) is a valid and sufficient scheme for labeling children's emotional expressions in this dataset.
- standard math Cohen's kappa and Jaccard similarity are appropriate metrics for evaluating speaker identification and emotion annotation agreement.
- domain assumption Manually labeled 'real speaker' references are ground truth for speaker identification.
- domain assumption Translated transcripts preserve emotional content for non-Finnish annotators.
- ad hoc to paper The session-count normalization for emotion distributions is a valid basis for comparing platforms.
Cite this review
Pith. "Pith review of Interaction Analysis by Humans and AI: A Comparative Perspective." pith.science (2026). https://pith.science/paper/IIYPRDQA
@misc{pith2026250607707,
author = {Pith},
title = {Pith review of: Interaction Analysis by Humans and AI: A Comparative Perspective},
year = {2026},
howpublished = {\url{https://pith.science/paper/IIYPRDQA}},
note = {Machine review of arXiv:2506.07707}
}
read the original abstract
This paper explores how Mixed Reality (MR) and 2D video conferencing influence children's communication during a gesture-based guessing game. Finnish-speaking participants engaged in a short collaborative task using two different setups: Microsoft HoloLens MR and Zoom. Audio-video recordings were transcribed and analyzed using Large Language Models (LLMs), enabling iterative correction, translation, and annotation. Despite limitations in annotations' accuracy and agreement, automated approaches significantly reduced processing time and allowed non-Finnish-speaking researchers to participate in data analysis. Evaluations highlight both the efficiency and constraints of LLM-based analyses for capturing children's interactions across these platforms. Initial findings indicate that MR fosters richer interaction, evidenced by higher emotional expression during annotation, and heightened engagement, while Zoom offers simplicity and accessibility. This study underscores the potential of MR to enhance collaborative learning experiences for children in distributed settings.
Figures
Reference graph
Works this paper leans on
-
[1]
Ruíz Gándara África, M Rosario González-Rodríguez, and M Carmen Díaz- Fernández. 2023. Salient features and emotions elicited from a virtual reality experience: the immersive Van Gogh exhibition.Quality & Quantity(2023), 1–20
work page 2023
-
[2]
Amazon Web Services. 2023. Amazon Transcribe Documentation. https://aws. amazon.com/transcribe/. Accessed: January 2025
work page 2023
-
[3]
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations.Advances in Neural Information Processing Systems (NeurIPS)33 (2020), 12449–12460. https://arxiv.org/abs/2006.11477
arXiv 2020
-
[4]
J. Cohen. 1960. A coefficient of agreement for nominal scales.Educational and Psychological Measurement20, 1 (1960), 37–46. https://doi.org/10.1177/ 001316446002000104
work page 1960
-
[5]
DeepL. 2023. DeepL Translator Documentation. https://www.deepl.com/. Ac- cessed: January 2025
work page 2023
-
[6]
Doccano Team. 2023. Doccano: Text Annotation Tool. https://github.com/ doccano/doccano. Accessed: January 2025
work page 2023
-
[7]
Naska Goagoses, Heike Winschiers-Theophilus, Jason Mendes, Selma Auala, and Erkki Sutinen. 2024. Primary School Students Designing for Future Relationship Building in Extended Realities: Encounters with Live Human Holograms. In Proceedings of the Participatory Design Conference 2024: Exploratory Papers and Workshops-Volume 2. 13–18
work page 2024
-
[8]
Google Cloud. 2023. Speech-to-Text Documentation. https://cloud.google.com/ speech-to-text. Accessed: January 2025
work page 2023
Show all 29 references
-
[9]
Alex Graves, Abdel rahman Mohamed, and Geoffrey Hinton. 2013. Speech recognition with deep recurrent neural networks. In2013 IEEE International Conference on Acoustics, Speech and Signal Processing. 6645–6649
2013
-
[10]
Sebastian Hahta, Maryam Teimouri, Tomi Suovuo, Selma Auala, Erkki Rötkönen, Jason Mendes, Naska Goagoses, Heike Winschiers-Theophilus, and Erkki Sutinen
-
[11]
Cheng Huang and Ming Dong. 2019. Text-based Emotion Detection using BERT and an Ensemble of Expert Features. InProceedings of the 2019 International Conference on Natural Language Processing and Knowledge Engineering. IEEE, 1–6
2019
-
[12]
Paul Jaccard. 1901. Étude comparative de la distribution florale dans une portion des Alpes et des Jura.Bulletin de la Société Vaudoise des Sciences Naturelles37 (1901), 547–579
1901
-
[13]
Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, et al
Melvin Johnson, Mike Schuster, Quoc V. Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, et al . 2017. Google’s multilingual neural machine translation system: Enabling zero-shot translation.Transactions of the Association for Computational Linguistics5 (2017), 339–351
2017
-
[14]
Leonid Keselman, John I Woodfill, Anders Grunnet-Jepsen, and Achintya Bhowmik. 2017. Intel RealSense Stereoscopic Depth Cameras. arXiv:1705.05548 [cs.CV] https://arxiv.org/abs/1705.05548
2017 arXiv
-
[15]
Kaitao Ma, Chunyang Xiao, and Jinho D. Choi. 2017. Text-based Speaker Identi- fication on Multiparty Dialogues Using Multi-document Convolutional Neural Networks. InProceedings of ACL 2017, Student Research Workshop. Association for Computational Linguistics, 49–55. https://ac...
2017
-
[16]
Paul Milgram and Fumio Kishino. 1994. A taxonomy of mixed reality visual displays.IEICE Transactions on Information and Systems77, 12 (1994), 1321–1329
1994
-
[17]
Anne Morris, Evelyne Maier, and Margaret Green. 2004. Character Error Rate: A New Evaluation Metric for Machine Translation.Computational Linguistics30, 2 (2004), 175–186
2004
-
[18]
Myriam Munezero, Calkin Suero Montero, Maxim Mozgovoy, and Erkki Sutinen
-
[19]
OpenAI. 2023. GPT-4 and GPT-3.5 Documentation. https://platform.openai.com/. Accessed: January 2025
2023
-
[20]
Sebeom Park, Shokhrukh Bokijonov, and Yosoon Choi. 2021. Review of Microsoft HoloLens Applications over the Past Five Years.Applied Sciences11, 16 (2021). https://doi.org/10.3390/app11167259
2021 doi
-
[21]
Robert Plutchik. 1980. A General Psychoevolutionary Theory of Emotion. In Theories of Emotion. Elsevier, 3–33. Interaction Analysis by Humans and AI: A Comparative Perspective IDC ’25, June 23–26, 2025, Reykjavik, Iceland
1980
-
[22]
René Riedl, Peter Mohr, Peter Kenning, Fred Davis, and Hauke Heekeren. 2011. Trusting humans and avatars: Behavioral and neural evidence. (2011). https: //aisel.aisnet.org/icis2011/proceedings/hci/7
2011
-
[23]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. InAdvances in Neural Information Processing Systems (NeurIPS), Vol. 30
2017
-
[24]
Kumar, Tim Fritsch, Marc Schröder, and Bjoern Schuller
Felix Weninger, Shakti P. Kumar, Tim Fritsch, Marc Schröder, and Bjoern Schuller
-
[25]
Yu Zhang, Jianfeng Yu, Hong Mei, and Bo Xu. 2019. Recurrent neural network- based error correction for Chinese automatic speech recognition.IEEE/ACM Transactions on Audio, Speech, and Language Processing27, 7 (2019), 1252–1265
2019
-
[26]
And indeed, I was afraid at first, but then I thought it might be fun, and I smiled. Still, I didn’t know what would happen next
Inc. Zoom Video Communications. 2023. Zoom Video Conferencing Software. https://zoom.us/. Accessed: January 2025. A Tables A.1 Character Error Rate (CER) Analysis This appendix provides detailed CER measurements across tran- scription tools, conversation types, and LLM-based c...
2023
-
[27]
arXiv:2309.05248 [cs.CL] https://arxiv.org/abs/2309
Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach. arXiv:2309.05248 [cs.CL] https://arxiv.org/abs/2309. 05248
-
[2013]
InProceedings of the 13th Koli Calling International Conference on Computing Education Research
Exploiting sentiment analysis to track emotions in students’ learning diaries. InProceedings of the 13th Koli Calling International Conference on Computing Education Research. 145–152
-
[2023]
In Supplementary Proceedings of the 11th International Conference on Communities & Technologies
Wearing a Hololens: A new dimension to remote presence in education. In Supplementary Proceedings of the 11th International Conference on Communities & Technologies. EUSSET. https://doi.org/10.48340/ct2023-2822
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.