{"id":"b3407f2d-599a-4e42-90ff-6e612443ac85","arxiv_id":"2607.28651","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Trained human coders agree more with each other than LLM labelers do on a 7-point ICAP cognitive-engagement scale, and human-refined coding rules do not transfer to LLM prompting.","lead":"This paper tested whether large language models can label how deeply students engage in group problem-solving as reliably as trained humans, using a 7-level version of the ICAP engagement framework. It found humans still agree with each other more than LLMs do, and giving LLMs the human-refined guidelines barely improved them.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline human-refinement gain (Δκ=0.10) is measured on the same two annotators who co-developed Criteria V2 and met regularly to resolve disagreements; without independent annotators, the improvement cannot be attributed to the framework.","rationale":"The reader's weakest_assumption identifies the same concern: human-human agreement gain is confounded by annotators co-developing the criteria and discussing disagreements. The paper itself concedes this in Limitations, but the abstract still presents the Δκ=0.10 as a headline finding. This is the most load-bearing issue because it undermines the central causal claim and also devalues the human ground truth used to evaluate LLM annotations. The paper's descriptive reliability numbers (human-human ~0.97, human-machine ~0.6, machine-machine ~0.85–0.89) are useful as a benchmark, and the honest Limitations section is creditworthy, but the headline causal claim should be presented with the confound clearly stated and ideally supported by independent-annotator evidence. The secondary issue about agent-refined frameworks is also valid but less fundamental. Given the reader already returned CONDITIONAL, our stress-test does not change that verdict; it reinforces the need for the proposed independent-annotator check before the causal claim can be accepted at face value.","tokens_in":11491,"tokens_out":2985,"duration_ms":27972,"concrete_test":"Recruit two or more naive annotators who were not involved in developing the criteria or in the original consensus meetings. Have them independently code a random sample of Stage 3 transcripts (or all transcripts) using Criteria V1 and V2 in a counterbalanced design with a washout period, with no inter-annotator discussion. Compute quadratic weighted kappa between the naive annotators for each version. If the independent Δκ is not significantly positive or is close to zero, the reported 0.10 gain is a co-development/discussion artifact rather than a framework property. Additionally, compare the naive annotators' labels to the original consensus labels to assess whether the gold standard is stable across annotator sets.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the human-refined framework improved human annotator agreement (Δκ=0.10) is load-bearing because it motivates the entire comparison and implicitly validates the human consensus as ground truth for evaluating LLMs. The data show H1H2 QWK rising from 0.906 in Stage 1 to 0.998 in Stage 3, but the same two annotators co-developed Criteria Version 2 in Stage 2→3, met regularly to resolve disagreements, and the first author mediated in Stage 1. The improvement may reflect shared learning, memory of discussed cases, and co-evolved interpretations rather than the written criteria alone. This confound inflates human-human reliability and also inflates any apparent success of the human-refined framework relative to LLM conditions, because the same annotators' consensus labels are used as ground truth for human-machine agreement. The paper's Limitations section explicitly concedes this: 'The high interrater reliability observed for human coders may partly reflect these opportunities for discussion' and recommends 'independent annotators outside the consensus process.' Yet the abstract still presents Δκ=0.10 as a headline result without this qualification. Without external validation, the causal attribution is not secure. A secondary issue is that the abstract's statement that 'Agent-refined frameworks improved cross-model agreement' is ambiguous: Table 2 shows agent-refined QWK (0.655–0.790) below human-refined QWK (0.874–0.894), so the claim only holds if the baseline is the original 4-point ICAP (0.613), not if compared to human-refined criteria.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper extends the ICAP framework to a 7-point ordinal scale for coding cognitive engagement in collaborative problem-solving discourse, and compares three annotation routes: trained human annotators, ICL/zero-shot LLM labeling with GPT-4o and GPT-5.2, and self-reflective LLM agents that iteratively refine the coding criteria. The authors report high human interrater reliability across three coding stages (QWK 0.906–0.998), moderate human–machine agreement (QWK ~0.59–0.66), and machine–machine agreement that is higher than human–machine but below human–human (QWK ~0.85–0.89 under human criteria, 0.66–0.79 under agent-refined criteria). The central claims are that the human-refined criteria improved human–human agreement by Δκ ≈ 0.10 but produced only small gains for ICL-based LLMs (Δκ < 0.04), and that agent-refined criteria improved cross-model agreement relative to the original 4-point ICAP but not to the level of human-refined criteria.","tokens_in":11706,"tokens_out":4383,"duration_ms":39804,"significance":"If the causal claims were secure, the paper would make a useful contribution to computational social science and educational data mining: it would show that theory-driven codebook refinement helps human annotators but does not transfer to black-box LLM prompting, and that self-reflective agents, while promising, do not yet match human-developed criteria for cross-model consistency. The paper has notable strengths: it makes code available, provides complete prompts and criteria versions in appendices, compares multiple chance-corrected agreement coefficients, and is unusually candid about procedural limitations. Those strengths are real. However, the headline causal claim about the human-refinement effect is confounded by shared annotator development and discussion, and the comparisons lack uncertainty quantification. The significance would be much higher if the human-refinement effect were demonstrated with independent annotators and with confidence intervals on all key kappa values.","major_comments":[{"comment":"The Δκ ≈ 0.10 improvement in human–human agreement (QWK rising from 0.906 in Stage 1 to 0.998 in Stage 3) cannot be causally attributed to the written criteria because the same two annotators co-developed Criteria Version 2, met regularly to resolve disagreements, and were mediated by the first author in Stage 1. Shared learning, memory of discussed cases, and co-evolved interpretations are plausible alternative explanations. The Limitations section acknowledges this ('may partly reflect these opportunities for discussion'), but the abstract presents Δκ = 0.10 as a headline result without this qualification. To make the claim load-bearing, the authors should either recruit independent annotators who did not participate in criteria development, or explicitly reframe the human-refinement result as exploratory and subject to this confound.","section":"Method: Human Annotation; Table 2"},{"comment":"All reported QWK, Krippendorff's α, and Randolph's κ values are point estimates without confidence intervals, significance tests, or multilevel/unit-clustered modeling. Many of the key comparisons—e.g., Δκ≈0.10 vs. Δκ<0.04, ICL vs. zero-shot differences of 0.01–0.04, and Stage 2 vs. Stage 3 differences—are of a magnitude that may fall within sampling variability, especially with N=42 group conversations and unequal stage sample sizes. The authors should report bootstrap or analytic CIs, and ideally model agreement at the group or trial level, before drawing conclusions about differential transfer of criteria to humans versus LLMs.","section":"Results: Interrater Reliability; Table 2"},{"comment":"The abstract's statement that 'Agent-refined frameworks improved cross-model agreement' is ambiguous. Table 2 shows agent-refined QWK values of 0.655–0.790, which are clearly below human-refined QWK values of 0.874–0.894. If 'improved' refers to improvement over the original 4-point ICAP (QWK=0.613 in Table 1), that baseline should be stated explicitly. Without specifying the comparison baseline, the abstract's phrasing may mislead readers into thinking agent refinement outperformed human refinement on cross-model consistency.","section":"Abstract; Table 2"},{"comment":"Human consensus labels serve both as the ground truth for constructing ICL examples and as the benchmark for evaluating LLM labels. While example transcripts are disjoint from test transcripts, the evaluation standard is the same consensus process that produced the examples. This creates a partial circularity, and it also means that any biases or idiosyncrasies in the human consensus labels are baked into both sides of the human–machine agreement comparison. The authors should at least acknowledge this explicitly and, ideally, evaluate on a held-out set of labels produced by independent annotators who were not part of the consensus discussions.","section":"Method: In-context Learning; Results: Human–Machine Reliability"}],"minor_comments":[{"comment":"The abstract uses 'kappa' generically, but Table 2 reports QWK, Randolph's κ, and Krippendorff's α. The Δκ = 0.10 value should specify which coefficient (appears to be QWK) and should be qualified as Stage 1 vs. Stage 3 of the same annotators.","section":"Abstract"},{"comment":"The table formatting is inconsistent: some cells contain semicolons, decimal alignment is uneven, and the notation 'H⋆M' is defined only in the note as 'agreement between the human consensus annotations and an LLM,' but the symbol '⋆' is not used consistently in the row labels. Also, the 'Cri-A1' and 'Cri-A2' rows under 'Zero-shot' appear to belong to M1M2 but their placement is easy to misread; consider merging or separating the sections more clearly.","section":"Table 2"},{"comment":"The right panel plots (1 − CosSim) between Criteria Version 1 and Version 2 for each engagement level, but the text does not define what CosSim measures (cosine similarity between which representations of the criteria?). A brief definition is needed for interpretability.","section":"Figure 3"},{"comment":"The paper does not state the number of ICL examples used per prompt, only that they were randomly sampled from a disjoint transcript. Since example count is a key hyperparameter that affects ICL behavior, please report it and, ideally, show a sensitivity analysis.","section":"Method: In-context Learning"},{"comment":"Criteria Version 2 labels Level 4 as 'Between Active and Constructive' while the main text refers to Level 4 as 'Initiating Constructive.' Please use consistent terminology across the main text and appendix.","section":"Appendix D"}],"recommendation":"major_revision","confidential_remarks":"The reader's conditional verdict is appropriate. The paper's main contribution is a valuable empirical comparison, but the central causal claim about human refinement is not yet secured because of the annotator confound, and the absence of uncertainty quantification weakens the comparative claims. This is fixable—by reanalyzing with independent annotators (or reframing), adding CIs, and clarifying the agent-refinement baseline—so I recommend major revision rather than rejection. The authors are unusually transparent about their limitations, which supports a constructive revision path."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague —\n\nQuick take: this is a reasonably honest, useful empirical paper, and the descriptive reliability numbers are probably worth keeping. But the headline result — that the human-refined framework raised human-human agreement by Δκ=0.10 — is not supported as a causal claim, because the same two annotators co-developed Criteria Version 2 and met regularly to resolve disagreements. The paper's Limitations section explicitly concedes this ('The high interrater reliability observed for human coders may partly reflect these opportunities for discussion'), yet the abstract still presents the gain without that qualification. That's the soft spot that matters most.\n\nWhat's genuinely new: the comparison of human annotation, ICL, zero-shot, and a self-reflective agent on a 7-point extension of ICAP, with consistency metrics across stages and criteria versions. The result that human-refined criteria improve human-human agreement but not ICL agreement (Δκ<0.04) is a useful negative finding for anyone building LLM annotation pipelines. The agent design — memory of recent samples and iterative criteria refinement — is a modest extension over MEGAnno+ and CrowdAgent, and the observation that agents only modified levels and never added/merged/deleted is an honest detail. The paper also credits Vosniadou et al. and prior annotation agents appropriately.\n\nSoft spots, in proportion: (1) the human-gain confound above, which also undermines the implicit validation of human consensus labels as ground truth for evaluating LLMs; (2) no confidence intervals or significance tests on any kappa comparisons, and labels are nested in 42 groups — a multilevel model or at least bootstrap intervals would be the obvious fix; (3) the abstract's claim that 'Agent-refined frameworks improved cross-model agreement' is ambiguous — it only holds if the baseline is the original 4-point ICAP (QWK 0.613), not if compared to human-refined criteria (0.87-0.89); (4) the agent memory buffer size (K=3) and number of ICL examples are reported in the algorithm but not ablated, and the code link has no commit hash or data release. These are mostly fixable with a revision.\n\nWho should read it: anyone working on LLM-based annotation of collaborative discourse or educational dialogue. It deserves a serious referee — the design flaws are real but the contribution is concrete and the limitations are partly acknowledged. I'd want an independent annotator replication and proper uncertainty quantification before trusting the headline.","headline":"Useful benchmark of human vs LLM coding on a 7-point ICAP scale, but the headline claim that human refinement improves agreement is confounded by the same annotators writing the criteria and discussing disagreements.","tokens_in":12348,"tokens_out":1875,"would_cite":true,"duration_ms":17243,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Refining a theory-derived coding framework boosts agreement between human annotators but barely changes what in-context-learning LLMs produce, and self-reflective agents remain below human-developed criteria.","keywords":["ICAP framework","7-point engagement scale","cognitive engagement","collaborative discourse","interrater reliability","in-context learning","self-reflective LLM agents","human-machine agreement"],"falsifier":"Code the same Stage 3 transcripts with fresh annotators who never participated in criteria development, using Version 1 and Version 2 in counterbalanced order; if their kappa gain is close to zero, the reported 0.10 gain is an artifact of shared training, not the framework.","tokens_in":11247,"feed_emoji":"🧠","tokens_out":8441,"duration_ms":71928,"temperature":0.7,"pith_summary":"This paper sets out to measure cognitive engagement in collaborative discourse and asks whether a refined coding framework transfers from human annotators to large language models. The authors built a 7-point engagement scale from the four ICAP modes, had two trained annotators apply it while they progressively sharpened the criteria, and compared the results with in-context learning, zero-shot prompting, and self-reflective LLM agents. The central finding is asymmetric transfer: the human-refined criteria raised human-human agreement by a kappa of about 0.10, whereas the same criteria moved LLM-based annotation by less than 0.04. Agent-driven refinement of the criteria produced better cross-model agreement than human-machine agreement but stayed below the level achieved when both models used human-developed criteria. The significance is practical: scalable automated annotation of engagement depends not only on clearer instructions but on the human process that produces them.","feed_headline":"Refined coding rules help humans, not LLMs","feed_subtitle":"A 7-point engagement scale lifted human-human kappa by 0.10 but moved in-context LLMs by under 0.04.","key_machinery":"The central object is an extended 7-point ICAP scale: an ordinal coding scheme that inserts intermediate levels between the four canonical engagement modes (Interactive, Constructive, Active, Passive) to capture gradations such as brief acknowledgements below active agreement or 'how/why' questions at the edge of constructive engagement. The argument runs through a refinement loop around this scale. Humans refine the criteria after disagreements; the authors give LLMs a parallel loop in which an agent labels a sample, reflects on its own uncertainty, and modifies, holds, or merges scale levels before proceeding. The scale gives both camps a common measurement language; the refinement loop is","core_discovery":"On the paper's own terms, the discovery is that the benefit of refining a coding framework depends on who does the refining and who applies it. Human annotators improved from quadratic weighted kappa (QWK) 0.906 at the first coding stage to 0.998 at the third after criteria were revised; the two large language models tested under in-context learning showed human-machine agreement that stayed near QWK 0.59–0.66 across criteria versions, with the refined version adding only 0.01–0.04. When two LLM agents were each allowed to reflect on ambiguous cases and update the criteria, agreement between the two models (QWK 0.655–0.790) was substantial but below the QWK 0.851–0.894 range that the same tw","pith_inferences":["A testable consequence the paper leaves implicit: because the same two annotators co-developed Criteria Version 2 and met regularly to resolve disagreements, part of the human kappa gain may be a shared-training effect; fresh annotators who never saw Version 1 would isolate the framework's contribution.","The agent chose only to modify existing levels, never to add, merge, or delete a level; this conservative behaviour may be an artifact of the reflection prompt, and a version that rewards structural changes might produce more human-like refinements.","If zero-shot prompting keeps matching or beating in-context learning as models improve, the marginal value of curated human exemplars may continue to fall, shifting the cost-benefit balance toward prompt design and evaluation rather than example curation.","A concrete way to test the multimodal gap is to give the LLM transcripts plus turn-level timing or audio cues from the same videos and check whether human-machine agreement moves toward the human-human level."],"forward_implications":["If the divergence is real, codebook refinement should be treated as a human-annotation activity; gains from clearer criteria will not automatically appear in LLM outputs when only prompt instructions are updated.","Human-consensus examples used in in-context learning do not reliably improve LLM-human agreement over zero-shot prompting, so few-shot example selection should not be assumed beneficial for theory-based coding tasks.","LLM agents can generate their own internally consistent criteria, but those criteria generalize less well across different foundation models than human-developed criteria do.","Mid-scale levels (3–5) are the main source of coding disagreement for both humans and agents, so future refinements should target decision rules at those boundaries.","Achieving scalable automated engagement coding may require fine-tuning on human-refined labels or multimodal input rather than better prompting alone."],"fun_headline_variants":["Coding rule updates help humans, barely move LLMs","Humans gain 0.10 kappa, LLMs gain under 0.04 from refined rules","Refined ICAP criteria: big human gains, tiny LLM gains","For coding engagement, humans benefit, LLMs don't"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing assumption is that the human consensus labels are the correct yardstick and that the rise in human-human agreement is caused by the refined criteria rather than by the two annotators having co-developed those criteria and repeatedly discussed disagreements.","fun_headline_variants_meta":{"raw":{"variants":["Coding rule updates help humans, barely move LLMs","Humans gain 0.10 kappa, LLMs gain under 0.04 from refined rules","Refined ICAP criteria: big human gains, tiny LLM gains","For coding engagement, humans benefit, LLMs don't"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1263,"prompt_tokens":776,"completion_tokens":487,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":407}},"tokens_in":520,"tokens_out":487,"duration_ms":4951,"temperature":1.0,"reasoning_tokens":407,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T00:47:54.536369+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Code the same Stage 3 transcripts with fresh annotators who never participated in criteria development, using Version 1 and Version 2 in counterbalanced order; if their kappa gain is close to zero, the reported 0.10 gain is an artifact of shared training, not the framework.","supporting_citations":[],"review_version":1}