{"id":"4ad56b2d-3655-4ead-bd71-e02cd8cb45a8","arxiv_id":"2507.01166","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A pilot study combines self-caught and probe-caught retrospective reporting to measure cognitive-affective states during collaborative learning, finding Optimistic, Curious, and Confused most frequent.","lead":"This paper tests a low-disruption way to capture how group members feel during collaboration: people watch a recording of their own session and click labels, either on their own or when prompted. If the method holds up, it could help build adaptive learning systems that respond to emotional and cognitive states in teamwork.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim rests on unvalidated temporal fidelity: probe-caught reports, collected on a fixed 60-second schedule, may reflect recall/response bias rather than in-the-moment states.","rationale":"The reader's verdict is CONDITIONAL, and my stress-testing converges on the same load-bearing assumption: temporal fidelity of retrospective self-report. The central claim of the paper is not that the method is fully validated, but that it can capture states with enough temporal detail to study dynamics. The results are descriptive counts and scatterplot trends. The most direct threat to this claim is that the measurement paradigm may systematically replace in-the-moment experience with post-hoc reconstruction. This is especially acute for probe-caught reports, since the probe schedule is arbitrary relative to the task. The authors themselves acknowledge the unresolved temporal fidelity in Section 6 and propose future multimodal validation. I also note a secondary internal ambiguity: Section 3.1.4 says the metadata includes 'report type' but then says reports were classified as self/probe by time differences between consecutive reports; this should be clarified, as it could affect the self/probe comparisons. However, the overall frequency and temporal patterns do not depend on that classification, so the core concern remains temporal fidelity. The concrete test I propose uses the already-collected multimodal data (video, audio, physiological) to provide convergent validation through coder annotations or automated signal detectors. This would settle whether the concern lands. Because the paper explicitly frames itself as a pilot and flags this limitation, the existing CONDITIONAL verdict is appropriate; no change is needed.","tokens_in":6671,"tokens_out":6064,"duration_ms":71246,"concrete_test":"Use the already-recorded Azure Kinect video, audio, and physiological data from the same 27 participants. Have two independent coders annotate the original task videos for visible confusion, frustration, and disengagement with event timestamps. Then compute the temporal alignment between coder-identified events and participant retrospective reports of the corresponding labels (e.g., percentage of retrospective reports falling within ±5 seconds of a coder-identified event, compared to a random-shift baseline). If alignment is at chance, the method's temporal fidelity is unsupported; if significantly above chance, the method gains convergent validity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that this retrospective cued-recall interface captures cognitive-affective states with enough temporal detail to study state dynamics—requires that each label and timestamp reported during video review corresponds to a genuine in-the-moment state experienced during the original task. This is not established. Section 6 explicitly concedes that 'the temporal fidelity of self-reports also remains undetermined.' The concern is not merely a missing validation; the measurement procedure structurally invites post-hoc reconstruction. Participants watch a video of their completed session and, especially in the probe-caught condition (Section 3.1.4), are interrupted every ~60 seconds and asked to report a state for that video moment. The fixed probe schedule is independent of actual state onset/offset, so a participant who cannot recall the exact moment may default to a generic label; the paper itself hypothesizes that the high rates of Optimistic and Curious are 'stand-ins for a neutral-engaged state' (Section 5). Thus the observed distributional differences between self-caught and probe-caught reports could reflect differential recall/response bias rather than actual state frequency or dynamics. Without an independent validation of temporal fidelity, the frequency counts (e.g., Optimistic=91) and temporal patterns (Curious early, Confused sustained, Disengaged rising) cannot be interpreted as evidence about real affective dynamics.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a pilot methodology for capturing cognitive-affective states of individuals in collaborative groups using retrospective cued recall. Participants first completed a collaborative weights task in groups of three; afterward they watched a video of their session and reported their internal states through an interactive interface. Reports were collected either as self-caught reports (participants voluntarily opened the survey) or probe-caught reports (the survey opened automatically after 60 seconds of idleness). The manuscript presents descriptive frequency statistics (359 labels total across 27 participants), compares the label distributions between self-caught and probe-caught reports, and describes temporal patterns in the reported states over normalized task time. The authors find that Optimistic, Curious, and Confused were the most frequent labels, and they interpret temporal patterns such as early Curious, sustained Confused, rising Disengaged, and mid-task Frustrated as consistent with prior work on affect dynamics. The paper explicitly frames itself as an initial analysis and acknowledges that the temporal fidelity of self-reports remains undetermined, calling for future validation against behavioral and physiological measurements.","tokens_in":6849,"tokens_out":3874,"duration_ms":164360,"significance":"If the reported method is validated, it would be a useful low-disruption tool for studying cognitive-affective state dynamics in collaborative learning, and it has clear relevance for educational data mining and adaptive learning systems. The paper benefits from a transparent statement of limitations, a reproducible description of the experimental procedure, direct measurement of report frequencies, and a publicly available source-code link. The descriptive observations about label frequencies and temporal distributions are concrete and checkable. However, the current manuscript does not yet establish the central validity claim that the captured reports faithfully reflect in-the-moment cognitive-affective states, and several quantitative comparisons are presented without appropriate statistical support. The paper is best read as a work-in-progress methodology; its significance as a published contribution depends on strengthening the link between the reported measures and actual state dynamics, or on carefully narrowing the claims to what the descriptive data can support.","major_comments":[{"comment":"The central interpretive claim that the frequency counts and temporal patterns describe actual cognitive-affective state dynamics relies on the assumption that retrospective cued recall is temporally faithful. Section 6 explicitly concedes that \"the temporal fidelity of self-reports also remains undetermined.\" With no validation against behavioral or physiological measurements, statements in Section 5 such as \"Confused reports persisted over longer periods without interruption\" and \"Frustrated reports peaked mid-task\" are not supported as claims about real state dynamics; they are observations about report timing. The manuscript should either add a validation component or systematically rephrase the Discussion and Conclusion to characterize the results as properties of the reporting behavior rather than of the underlying affective states.","section":"§5, §6"},{"comment":"The comparison between self-caught and probe-caught distributions is based on unequal totals (129 versus 230 labels) and is not accompanied by any statistical test or normalized rate. Because participants could report multiple labels per survey, raw label counts conflate the number of reports with the number of labels per report. The claim that the distributions \"changed across the two collections\" is therefore not substantiated. The authors should report per-report proportions or rates, and apply an appropriate inferential test (or explicitly state that all conclusions are purely descriptive and forgo comparative claims).","section":"§4, Table 2, Figures 4–5"},{"comment":"The classification of reports as self-caught versus probe-caught is based on the time difference between consecutive reports \"align[ing] exactly with the probe-frequency.\" This rule is underspecified: no tolerance or jitter is defined, the 60-second probe interval interacts with the 1-second video skip after self-reports, and survey completion time could shift the next report timestamp. Without a precise, robust decision rule and a sensitivity check, misclassification between the two conditions could bias the distributional comparisons in Table 2 and Figures 4–5.","section":"§3.1.4"},{"comment":"The \"Other\" response option was added after the second experimental group, so groups 1–2 had a different, smaller label set than groups 3–9. Pooling all participants for the frequency distributions and self- versus probe-caught comparisons therefore mixes two different measurement conditions. At a minimum, the authors should report whether the main frequency patterns hold when restricted to the subset of participants who had the full label set, and should note this inconsistency explicitly in the Results.","section":"§3.1.4, §3.2, §4"}],"minor_comments":[{"comment":"The text cites Khebour et al. and the task-generalization concern with reference [20], but reference [20] in the bibliography is Highhouse's \"Designing experiments that generalize\"; the in-text citation numbering appears to be shifted and should be corrected throughout.","section":"§3.1.2 and §6"},{"comment":"Table 3 reports an average inter-report time difference of 40.43 seconds, which is notably shorter than the 60-second probe interval; the authors should explain this discrepancy, for example by separating self-caught and probe-caught intervals or by reporting the effect of the 1-second video skip and multiple self-reports.","section":"§4, Table 3"},{"comment":"The row labels in Table 2, particularly \"Mean,\" \"SD,\" and \"Standard Error,\" should specify the unit of analysis (labels per participant or per report) so that the descriptive statistics are unambiguous.","section":"§4, Table 2"},{"comment":"The statement that participants \"reported more low-arousal emotions when probed\" is not tied to a definition of low-arousal or to a quantitative comparison; either provide the supporting analysis or remove the claim.","section":"§5"},{"comment":"The scatterplot in Figure 6 would be easier to interpret if the axes were labeled directly in the figure and if overplotted points were given some visual transparency or jitter, since many reports occur near the same normalized times.","section":"§4, Figure 6"}],"recommendation":"major_revision","confidential_remarks":"This is a pilot methodology paper with an honest limitations section. The central issue is that the manuscript's interpretive claims extend beyond what the descriptive data and unvalidated retrospective recall procedure can support. I believe the paper is salvageable with substantial revision: either add a credibility check for temporal fidelity (even a small one) or systematically soften the language about state dynamics, and add appropriate statistical treatment of the self- versus probe-caught comparisons. I also note that the reference list has at least one clear citation-numbering error, which suggests the manuscript needs careful proofreading before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a small, honest pilot study that extends the authors' own retrospective cued-recall paradigm with a self-caught versus probe-caught structured survey. The new dataset and the two-mode design are real additions, and the paper is refreshingly candid about its limits. But the load-bearing assumption—that retrospective reports preserve in-the-moment temporal fidelity—is not validated, and the authors concede exactly this in Section 6. So the frequency counts and temporal patterns should be read as descriptive outputs of the method, not as evidence about actual affective dynamics.\n\nWhat's genuinely new: the structured survey with a fixed 60-second probe schedule plus voluntary self-caught reports, applied to collaborative problem-solving groups. The 27-participant corpus, with time-aligned labels and scatterplot visualizations, is new. The simple classification rule for self- versus probe-caught based on inter-report intervals is a neat operational detail, though it depends on exact alignment.\n\nWhat the paper does well: it is clearly written, the procedure is reproducible from the description, and the limitations discussion is honest—temporal fidelity, small n, generalizability, and the possibility that Optimistic and Curious are stand-ins for a neutral-engaged state are all flagged. It does not oversell the results.\n\nThe soft spots are real but proportionate to the pilot status. First, temporal fidelity is the crux, and Section 6 leaves it open; the stress-test concern about the 60-second probe schedule is valid, because the probes are tied to clock time, not to state onsets or behavioral cues. Second, the comparison between self-caught (129 labels) and probe-caught (230 labels) mixes different denominators, and no statistical test accompanies the claimed distributional differences. Third, the temporal narrative about Curious early, Confused sustained, Disengaged late, and Frustrated mid-task is based on eyeballing a scatterplot, not any test. These are limitations, not fatal flaws, given the paper frames itself as an initial analysis.\n\nThe citation pattern is fine; it builds on the authors' prior work and D'Mello and Graesser, and the self-reference is legitimate. No equation-level circularity.\n\nWho should read it: people working on affect measurement methods in educational data mining, especially those interested in retrospective recall and low-disruption annotation. It doesn't deserve a desk reject; a serious referee could push for validation against behavioral or physiological data, and for statistical rigor in the comparisons.\n\nMy recommendation: send it out for review with a clear request that the authors either add validation or sharpen the claims to match the evidence.","headline":"Honest pilot study with a useful new data collection variant, but the temporal fidelity assumption is unvalidated—so treat the frequency and temporal claims as descriptive, not confirmatory.","tokens_in":7438,"tokens_out":2304,"would_cite":false,"duration_ms":25608,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a retrospective cued recall procedure can timestamp cognitive-affective states in collaborative groups without interrupting the task itself.","keywords":["Collaborative Learning","Cognitive-Affective States","Retrospective Cued Recall","Self-Caught Reports","Probe-Caught Reports","Affective Computing","Educational Data Mining","Temporal Dynamics"],"falsifier":"Run the same cued recall procedure alongside concurrent behavioral and physiological recording, then align the reported onset times for states like Confused and Frustrated with machine-coded facial action units, posture shifts, or voice cues from the original session. If the report timestamps consistently lag or lead the marker signals by large, variable amounts, the method's temporal fidelity claim fails; if they align within a small window, the method's core assumption is supported.","tokens_in":6421,"feed_emoji":"🧠","tokens_out":8636,"duration_ms":91765,"temperature":0.7,"pith_summary":"This paper claims that a retrospective cued recall procedure—participants watch a video of their own collaborative session and report cognitive-affective states either on their own or when prompted—can capture those states with temporal detail without disrupting the collaboration itself. Drawing on 27 participants in nine three-person groups solving a weights task, the paper reports an initial analysis of 359 labeled reports in which Optimistic, Curious, and Confused dominate. The temporal scatterplot shows early curiosity, sustained confusion, rising disengagement, and mid-task frustration, and the authors read these patterns as consistent with existing accounts of affective dynamics. If the method is sound, it would give the educational data mining community a low-disruption measurement tool for studying how group-level cognitive-affective states unfold and for building adaptive learning systems that react to them.","feed_headline":"Video recall captures group emotions without interrupting teamwork","feed_subtitle":"Participants timestamp confusion and curiosity while rewatching their session, giving adaptive systems data to act on.","key_machinery":"The central object is a retrospective cued recall interface: an interactive application plays back the group's session video and opens a survey either when the participant chooses to report (self-caught) or automatically after about 60 seconds of idleness (probe-caught). The survey constrains reports to a pre-identified set of cognitive-affective states—Confused, Disengaged, Curious, Optimistic, Frustrated, Conflicted, Surprised, plus Other—and lets participants report multiple states at once with an onset/ongoing distinction. The mechanism for temporal analysis is the report timestamp: timestamps are normalized within each group by the group's video length, which turns the raw reports into a scatterplot of label occurrences over task time. Frequency distributions and descriptive statistics of self-caught versus probe-caught reports are then compared as a check on whether the two collection modes yield similar pictures.","core_discovery":"The central claim is that a structured, video-stimulated recall survey can produce a usable timestamped trace of the cognitive-affective states individuals experience during collaborative problem solving. In the reported data, participants used a fixed menu of seven labels plus an 'Other' option, reported multiple overlapping states freely, and marked each state as onset or ongoing; the system recorded either self-caught reports or probe-caught reports triggered after 60 seconds without input. The resulting 359 labels were led by Optimistic (91), Curious (82), and Confused (66), with self-caught and probe-caught subsets showing a similar top three. Over normalized task time, Curious appeared early, Confused persisted across long stretches, Disengaged rose in the latter half, and Frustrated peaked mid-task before dropping off; the paper takes these temporal shapes as evidence that the paradigm captures state dynamics, not just aggregate frequencies.","pith_inferences":["Because the menu is fixed to seven labels, the high counts for Optimistic and Curious may partly reflect label availability rather than true prevalence; adding an explicit 'neutral-engaged' label, which the paper suggests participants were reaching for, could redistribute the frequencies.","The 60-second probe interval sets a practical floor on how short a state must be to appear in probe-caught data, so very brief affective events may be underrepresented in this collection mode.","A direct test of the method's temporal fidelity is available in the existing recordings: align report timestamps with facial action units or physiological signals from the same session and measure the lag between marker onset and report time.","If the approach transfers to other tasks, the normalized-time scatterplot could become a standard diagnostic for comparing state dynamics across different collaborative learning activities."],"forward_implications":["If the method is valid, cognitive-affective states in collaborative groups can be timestamped without interrupting the task itself, since reports happen during video review rather than during collaboration.","The similar top-three label distributions under self-caught and probe-caught collection suggest that a 60-second auto-probe can broaden coverage without heavily distorting which states people report.","The temporal patterns—early Curious, sustained Confused, rising Disengaged, mid-task Frustrated—offer concrete hypotheses about collaborative state dynamics that future studies with larger samples can test.","Because each report carries a timestamp, the data can be aligned with behavioral and physiological recordings, enabling multimodal models that connect reported states to observable cues, which the paper names as future work.","Affect-aware adaptive learning systems could use such timestamped reports to detect when a group is persistently confused or disengaging and respond in time."],"supporting_citations":[{"why":"Earlier retrospective cued recall study that identified the cognitive-affective label themes this survey formalizes.","marker":"[1]"},{"why":"Precedent for learners watching and judging their own session after a task, the core mechanism of cued recall.","marker":"[12]"},{"why":"Theory used to interpret Optimistic and Curious as stand-ins for a neutral-engaged state and frustration as an intermediate response.","marker":"[14]"},{"why":"Supports reading Confused as a sustained state in the temporal analysis.","marker":"[15]"},{"why":"Supports treating Surprised as intermittent and transitory in the temporal analysis.","marker":"[3]"},{"why":"Documents the recall-bias and temporal-resolution problems of post-task self-reports that the method is designed to mitigate.","marker":"[24]"}],"fun_headline_variants":["Video recall timestamps group emotions without interrupting flow","Probe-caught reports reveal state dynamics in teamwork","Curious early, frustrated mid-task: video recall maps affects","Self-report vs probe: capturing affect in collaborative settings","Video recall data could adapt learning systems to team states"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Participants can accurately recall and timestamp the cognitive-affective states they felt during the task while watching their own video, rather than reconstructing them after the fact.","fun_headline_variants_meta":{"raw":{"variants":["Video recall timestamps group emotions without interrupting flow","Probe-caught reports reveal state dynamics in teamwork","Curious early, frustrated mid-task: video recall maps affects","Self-report vs probe: capturing affect in collaborative settings","Video recall data could adapt learning systems to team states"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000554,"raw_usage":{"total_tokens":2591,"prompt_tokens":852,"completion_tokens":1739,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":468,"completion_tokens_details":{"reasoning_tokens":1661}},"tokens_in":468,"tokens_out":1739,"duration_ms":17313,"temperature":1.0,"reasoning_tokens":1661,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:58:28.615451+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same cued recall procedure alongside concurrent behavioral and physiological recording, then align the reported onset times for states like Confused and Frustrated with machine-coded facial action units, posture shifts, or voice cues from the original session. If the report timestamps consistently lag or lead the marker signals by large, variable amounts, the method's temporal fidelity claim fails; if they align within a small window, the method's core assumption is supported.","supporting_citations":[{"cited_title":"Here, we present initial experiments with retrospective cued recall paradigm for identifying cognitive-affective states of individ- uals in collaborative groups","cited_arxiv_id":null,"evidence_quote":"Earlier retrospective cued recall study that identified the cognitive-affective label themes this survey formalizes."},{"cited_title":"Bosch, S","cited_arxiv_id":null,"evidence_quote":"Precedent for learners watching and judging their own session after a task, the core mechanism of cued recall."},{"cited_title":"Surprised","cited_arxiv_id":null,"evidence_quote":"Theory used to interpret Optimistic and Curious as stand-ins for a neutral-engaged state and frustration as an intermediate response."},{"cited_title":"Chandler, T","cited_arxiv_id":null,"evidence_quote":"Supports reading Confused as a sustained state in the temporal analysis."},{"cited_title":"Confused","cited_arxiv_id":null,"evidence_quote":"Supports treating Surprised as intermittent and transitory in the temporal analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the recall-bias and temporal-resolution problems of post-task self-reports that the method is designed to mitigate."}],"review_version":1}