{"id":"a85e3670-2bba-415a-b1fb-a4d40c44b367","arxiv_id":"2507.02229","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In a Lego-based collaborative task, participants' retrospective verbal reports show distinct linguistic patterns, with positive emotion labels such as 'Engaged' and 'Optimistic' appearing most frequently.","lead":"Researchers had 32 people build a Lego structure in teams of four, then each person watched a replay of the task and spoke aloud what they had been thinking and feeling. The study found distinct language patterns in these monologues, including frequent emotion words, and proposed a new set of internal-state emotion categories for collaborative work.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Label circularity is the load-bearing weakness: the 61.4% positive-prevalence figure depends on an unreported, non-exclusive mapping of overlapping keywords to labels, so the distribution claim is not yet supported.","rationale":"The reader correctly flags retrospective recall as a validity threat, and I agree it is real and acknowledged in Section 6. But I do not think it is the single most load-bearing point, because the paper's quantitative centerpiece—the 61.4% positive-label prevalence—would be uninterpretable even if recall were perfect. Table 8 assigns the same surface keywords to multiple labels (e.g., 'figure out' to both Engaged and Optimistic), and no inter-rater reliability is reported. The semantic-similarity check in Section 4.4 does not break this circularity: grouping words by semantic relatedness and then measuring semantic relatedness within the groups is close to a tautology. The authors are transparent about the exploratory status and limitations, and the n-gram observations (e.g., 'confus' being most frequent) are useful descriptive results. However, the central prevalence claim needs an independent annotation pass or a clearly exclusive coding rule before it can be treated as evidence about which internal states are most or least prevalent. This keeps the verdict at CONDITIONAL; it does not reject the paper's qualitative contributions.","tokens_in":8988,"tokens_out":5543,"duration_ms":66011,"concrete_test":"Release the de-identified transcripts and have two independent annotators assign every emotion-relevant keyword occurrence to exactly one of the twelve labels (with an 'other' option), following a pre-registered codebook. Compute Cohen's kappa; then recompute Figure 2 and the 61.4% positive share using only mutually exclusive assignments. If kappa is below about 0.6, or if the positive share shifts materially (e.g., below 50%), the prevalence claim in Section 4.3 is an artifact of the original non-exclusive mapping. This single reanalysis uses data already collected and does not require new experiments.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing threat to the central claim is not only retrospective recall but the internal validity of the emotion-label counts. The twelve labels in Table 8 were created post hoc from the same keyword list used to measure prevalence, and no rule is reported for assigning each keyword occurrence to exactly one label. Several surface forms are listed under multiple labels: 'figure out' appears under Engaged and Optimistic; 'easier'/'easi' under Optimistic and Satisfied; 'vagu' under Conflicted and Confused; 'assum' under Conflicted and Reserved; 'mistak' under Anxious and Disappointed; 'crazi' under Frustrated and Anxious. If occurrences are counted once per label, positive labels with more assigned keywords are systematically inflated, making the headline 61.4% in Section 4.3 an artifact of the uncontrolled mapping rather than a measured prevalence. The Section 4.4 semantic-similarity validation is also circular: words were grouped because they are semantically related, so high BERT cosine similarity within categories is expected, and no random baseline or inter-rater reliability is reported. Concretely, a second coder re-applying the codebook, or a context-based resolution of overlapping keywords into a single label, could substantially change the reported distribution. The exploratory n-gram observations are credible; the quantitative prevalence claim is not yet supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an exploratory linguistic analysis of retrospective self-reports collected from 29 participants immediately after a four-person Lego collaborative problem-solving task. Participants watched a video replay of the task and narrated their moment-to-moment internal states; the authors transcribed these monologues, applied frequency analyses of unigrams, bigrams, and trigrams, hand-selected 29 keywords, grouped them into twelve emotion labels, and report that \"Optimistic,\" \"Engaged,\" and \"Satisfied\" account for 333 (61.4%) of emotion-label word occurrences. A BERT-based semantic-similarity analysis (within- and between-category cosine similarities) is presented as validation that the labels are coherent and distinct.","tokens_in":9253,"tokens_out":6003,"duration_ms":61922,"significance":"The study addresses a real gap: internal states during CPS are usually inferred from observable behavior, and retrospective replay is a relatively underexplored method for accessing individual experience. The paper is commendable for making its full keyword tables and label mappings visible, for transparently reporting demographic and exclusion information, and for acknowledging the retrospective-recall limitation in Section 6. If the label-based prevalence claims were supported by a reproducible matching rule and independent validation, the dataset would be a useful descriptive baseline for future CPS emotion research. As it stands, the exploratory n-gram findings are credible, but the headline distributional claim is not yet supported.","major_comments":[{"comment":"The 333-occurrence/61.4% figure is not well-defined because Table 8 lists the same surface keywords under multiple labels without a stated counting rule. For instance, \"figure out\" appears under Engaged and Optimistic; \"easier\"/\"easi\" under Optimistic and Satisfied; \"vagu\" under Conflicted and Confused; \"assum\" under Conflicted and Reserved; \"mistak\" under Anxious and Disappointed; and \"crazi\" under Frustrated and Anxious. If each occurrence is counted once per label, positive labels with duplicated keywords are systematically inflated; if occurrences are assigned to one label only, the adjudication rule is missing. Table 8 also contains words not in Table 7 (e.g., \"zoning out,\" \"conflict,\" \"sure,\" \"stress,\" \"worri,\" \"fault,\" \"lost,\" \"surpris\"), so the source of the occurrence counts in Figure 2 is unclear. Please report a complete, single-label mapping from each keyword/phrase occurrence to one label, or a precise multi-label counting rule, and provide a sensitivity analysis of the 61.4% figure under alternative assignments.","section":"Section 4.3, Table 8, Figure 2"},{"comment":"The semantic-similarity validation is circular. The labels and their member words were constructed post hoc from the same keyword-frequency data by the same authors, so high within-category BERT cosine similarity is expected and does not independently confirm that the categories are coherent. No random baseline, shuffled-label comparison, or inter-rater reliability is reported. Please add a baseline such as average similarity of random word sets of matched size, a second coder re-applying the codebook to transcripts, or a hold-out coding exercise, so that Table 9 can be interpreted as evidence for label coherence.","section":"Section 4.4, Table 9"},{"comment":"The entire dataset consists of retrospective verbal reports produced during video replay, and the paper itself states in Section 6 that \"participants may forget details or not accurately report their emotional experiences after the task is completed.\" Because the paper's central claims concern internal states during the task, the absence of any concurrent or independent check on the fidelity of recall leaves the construct validity of the frequency tables open. Either provide evidence for the fidelity of the replay procedure (e.g., alignment with known task events, or a subset with in-task measures), or revise the wording throughout Sections 4 and 5 so that the results are explicitly framed as characterizations of retrospective accounts rather than directly measured in-task states.","section":"Sections 3.1 and 6"}],"minor_comments":[{"comment":"\"When mapping, we accounted for the the context used by participants\" contains a duplicated \"the.\"","section":"Section 3.3"},{"comment":"The frequency cutoffs for unigrams (>=30), bigrams (>=6), and key terms (>=5) are introduced without a rationale; a sentence on why these thresholds were chosen would help.","section":"Section 3.3"},{"comment":"The text says the age range was 20 to 32, while the table groups ages as 18-24, 25-31, and 32+; please reconcile these descriptions.","section":"Section 3.1, Table 1"},{"comment":"The trigram analysis reports phrases with frequencies between 3 and 5, but no trigram cutoff is specified in Section 3.3; please state the criterion used.","section":"Section 4.1"},{"comment":"The discussion of highest and lowest between-category similarities would be easier to check if the specific numeric values or a full similarity matrix were provided in the text.","section":"Section 4.4, Figure 4"},{"comment":"The paper does not include a data or code availability statement; if the de-identified transcripts can be shared, adding such a statement would substantially improve reproducibility and would directly help reviewers evaluate the label-mapping concerns.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest and exploratory, and the central weaknesses are methodological rather than malicious. I believe the label-mapping and validation issues can be addressed within the scope of a revision; the current form is not yet ready for publication because the headline 61.4% prevalence figure depends on an unreported counting rule and the semantic-similarity validation is not independent."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely useful exploratory dataset wrapped in a too-strong interpretive frame. I agree with the conditional verdict, and the stress-test note on label circularity lands.\n\nWhat's new: a new CPS task with 8 four-person teams, 29 usable retrospective monologues collected by having participants watch the GoPro video of the first phase and speak their moment-to-moment thoughts. The raw n-gram tables are plausible, and the twelve-category emotion label set (Engaged, Optimistic, Satisfied, Confused, etc.) is more granular than what the cited CPS literature uses. The authors are transparent about the obvious limits: small sample, self-report bias, excluding three builder recordings, and no role-level analysis. That transparency makes the paper easy to read and build on.\n\nThe soft spots are the same ones the stress-test flags. The central prevalence claim — three positive labels account for 61.4% of emotion words — depends on a post hoc label-to-keyword mapping (Table 8) that is never fully specified. Several surface forms appear under multiple labels: 'figure out' in Engaged and Optimistic, 'easi' in Optimistic and Satisfied, 'vagu' in Conflicted and Confused, 'mistak' in Anxious and Disappointed, 'crazi' in Frustrated and Anxious. The paper does not say how each occurrence was assigned to one label, or whether occurrences were counted under every matching label. Without that, the frequency distribution is not a measured prevalence; it is an artifact of the mapping decisions. The same issue shadows the BERT similarity 'validation': words were placed together because they are semantically close, so high within-category cosine similarity is expected and proves little. No random baseline and no inter-rater agreement are reported.\n\nThis is fixable. The authors should release the transcripts, add a second annotator, report Cohen's kappa, resolve overlapping keywords with a written rule, and re-frame Section 4.4 as descriptive rather than confirmatory. They should also either include the remaining builder recordings or explicitly state the role imbalance as a scope limit — they do mention it, but the exclusion should be justified more carefully.\n\nWho should read it: people designing emotion annotation schemes for CPS or studying self-report methodology. It is not a definitive empirical result, but it is a legitimate foundation. I would send it to review, with the expectation of major revision. The data collection design is reproducible and the research question is worth pursuing.","headline":"Honest, useful exploratory dataset, but the 61.4% positive-prevalence headline is not yet supported: the label mapping is post hoc, overlapping, and unreported.","tokens_in":9799,"tokens_out":2575,"would_cite":false,"duration_ms":29204,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Retrospective video-recall monologues reveal a twelve-label map of internal states during collaborative problem solving.","keywords":["collaborative problem solving","internal states","emotion labels","retrospective self-report","stimulated recall","linguistic analysis","semantic similarity","n-gram frequency"],"falsifier":"Collect concurrent in-task measures of emotion, such as physiological arousal or button-press emotion ratings, during the collaborative task and compare them with the retrospective label distribution from the video-recall monologues; if the two show little or no correspondence, the claim that retrospective speech reveals actual internal states would be falsified.","tokens_in":8784,"feed_emoji":"🧠","tokens_out":5499,"duration_ms":60558,"temperature":0.7,"pith_summary":"This paper argues that the internal emotional states people experience while collaborating can be recovered from their own spoken reflections, collected immediately after the task as they watch a video replay of the team working. Analyzing 29 participants' monologues from a four-person Lego building task, the authors identify frequent words and phrases, group them into twelve emotion labels, and report the distribution of those labels. Three positive labels—Optimistic, Engaged, and Satisfied—account for 333 (61.4%) of emotion-label word occurrences, while Disengaged, Reserved, and Frustrated together make up only 24 (4.4%). The value of the claim, if it holds, is a cheap and scalable way to study individual internal states in collaborative settings without interrupting the team's work.","feed_headline":"Positive internal states take 61% of emotion words in teamwork","feed_subtitle":"Analyzing 29 participants' video-recall monologues, researchers mapped collaborative problem solving onto twelve emotion labels.","key_machinery":"The load-bearing mechanism is the stimulated-recall protocol: each participant, alone after the task, watches a video of their team and narrates their thoughts and feelings moment-to-moment, with the video allowing those states to be mapped back to task time. On top of that transcript, the analysis builds a small emotion taxonomy by hand: 29 frequent keywords and phrases are assigned to twelve labels, and the labels are validated by within-category and between-category cosine similarities computed from pretrained transformer word embeddings. The taxonomy is what turns free-form monologue into a quantitative distribution of internal states.","core_discovery":"The central claim is that internal states during collaborative problem solving leave reliable traces in retrospective speech, and that a hand-derived set of twelve emotion labels—Engaged, Disengaged, Conflicted, Confident, Reserved, Frustrated, Optimistic, Anxious, Disappointed, Satisfied, Confused, Surprised—captures the distribution of those states. The authors show that frequent n-grams are mostly filler, so they manually select 29 content-bearing keywords and phrases, assign them to labels using discrete-emotion literature, and then use cosine similarity of word embeddings to show within-label words are semantically close and labels are reasonably distinct. The resulting frequency distribution is skewed positive: engagement, optimism, and satisfaction dominate, while frustration, disengagement, and reservedness are rare.","pith_inferences":["Inference: if the retrospective protocol is validated against concurrent measures, the same method could be extended to role-level analysis, since directors and builders may show different emotion profiles.","Inference: the label set is likely task- and context-specific; the Lego construction task may induce more positive states than competitive or high-stakes collaboration, so the 61% positive figure should not be generalized without replication.","Inference: the semantic-similarity validation could be turned into a testable prediction that automatic classifiers trained on the twelve labels should predict task-relevant events, such as a builder making an error or a director giving confusing instructions, better than chance.","Inference: one could test whether the act of narrating changes the experience by comparing first-phase monologues with second-phase task behavior, a check for reactivity of the method."],"forward_implications":["Future collaborative problem solving studies can use prompted video recall to collect internal-state data at scale without interrupting collaboration.","The twelve labels and their associated keywords give later researchers a starting vocabulary for automatic emotion annotation.","The distribution suggests positive states dominate in this kind of cooperative construction task, so interventions aimed at reducing frustration may target a minority of experience.","Because confusion is relatively frequent and known to accompany learning, its presence can be treated as a sign of engagement rather than failure.","The within-category and between-category semantic similarity method offers a template for checking that hand-built emotion categories are coherent."],"supporting_citations":[{"why":"Supplies the discrete-emotion theory that grounds the twelve hand-derived emotion labels.","marker":"[12]"},{"why":"Provides the speech-recognition model used to produce initial transcripts of the self-report recordings.","marker":"[33]"},{"why":"Provides the stop-word and punctuation removal used in preprocessing transcripts.","marker":"[27]"},{"why":"Supplies the within-category versus between-category comparison design used to validate the emotion taxonomy.","marker":"[18]"},{"why":"Defines cosine similarity, the metric used to compute semantic closeness of word embeddings.","marker":"[34]"},{"why":"Supplies the bidirectional transformer language model used to produce word embeddings for semantic similarity.","marker":"[23]"},{"why":"Describes the fine-tuned uncased variant of that language model used in the actual embedding computation.","marker":"[16]"},{"why":"Establishes the collaborative laboratory task paradigm that the Lego building exercise follows.","marker":"[2]"},{"why":"Supports the interpretation that confusion can be beneficial for learning, used in discussing the Confused label.","marker":"[11]"}],"fun_headline_variants":["Teamwork emotions lean positive in recall","Positive states dominate collaborative problem solving","Collaborators recall more positive than negative emotions","Emotion labels in teamwork skewed positive"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire dataset is retrospective speech, so the paper must assume that what participants say while watching the video accurately reconstructs what they actually felt during the task; if memory or social desirability distorts those recollections, the emotion-label frequencies describe the retelling rather than the experience.","fun_headline_variants_meta":{"raw":{"variants":["Teamwork emotions lean positive in recall","Positive states dominate collaborative problem solving","Collaborators recall more positive than negative emotions","Emotion labels in teamwork skewed positive"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000302,"raw_usage":{"total_tokens":1670,"prompt_tokens":809,"completion_tokens":861,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":425,"completion_tokens_details":{"reasoning_tokens":810}},"tokens_in":425,"tokens_out":861,"duration_ms":10321,"temperature":1.0,"reasoning_tokens":810,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:33:52.638858+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect concurrent in-task measures of emotion, such as physiological arousal or button-press emotion ratings, during the collaborative task and compare them with the retrospective label distribution from the video-recall monologues; if the two show little or no correspondence, the claim that retrospective speech reveals actual internal states would be falsified.","supporting_citations":[{"cited_title":"In: International conference on machine learning","cited_arxiv_id":null,"evidence_quote":"Provides the speech-recognition model used to produce initial transcripts of the self-report recordings."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the within-category versus between-category comparison design used to validate the emotion taxonomy."},{"cited_title":"In: The 7th international student conference on advanced science and technology ICAST","cited_arxiv_id":null,"evidence_quote":"Defines cosine similarity, the metric used to compute semantic closeness of word embeddings."},{"cited_title":"International Journal of Intel- ligent Networks2, 64–69 (2021)","cited_arxiv_id":null,"evidence_quote":"Describes the fine-tuned uncased variant of that language model used in the actual embedding computation."},{"cited_title":"Science329(5995), 1081–1085 (2010)","cited_arxiv_id":null,"evidence_quote":"Establishes the collaborative laboratory task paradigm that the Lego building exercise follows."},{"cited_title":"Learning and Instruction29, 153–170 (2014)","cited_arxiv_id":null,"evidence_quote":"Supports the interpretation that confusion can be beneficial for learning, used in discussing the Confused label."}],"review_version":1}