{"id":"1e2ddda5-c0fc-4174-81f7-856012f11cff","arxiv_id":"2607.17452","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Acoustic feature reliability in dyadic remote interaction largely reflects speaker identity, linguistic features predict well but don't generalize across tasks, and only interaction features survive speaker normalization.","lead":"This paper tests whether acoustic, linguistic, and interaction features reliably measure cognitive load and power in remote dyadic conversations, and finds that acoustic \"reliability\" mostly reflects speaker identity. It proposes a three-way evaluation—prediction, cross-task generalization, and test-retest reliability—as a standard for choosing multimodal features.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Acoustic-ICC collapse may be driven by unmatched test-retest demand levels rather than speaker identity.","rationale":"The paper is a careful empirical study, and I read it in good faith: the three-dimensional framework is sensible, the LODO/LOCO/ICC distinction is appropriate, and the reported numbers are internally consistent. The reader's conditional verdict is justified. However, the single most load-bearing point in the paper is the acoustic reliability collapse, because the conclusion 'commonly used prosodic features measure who speaks rather than how speakers behave under load' generalizes a warning to the whole multimodal affective computing community. The reader's identified weakness (speaker-specific task responses being erased by z-scoring) is closely related but not the sharpest threat: within-speaker z-scoring actually preserves a speaker's consistent directional response to load, so that version of the concern is weaker than stated. The sharper threat is that the two test-retest tasks themselves are not matched on the state variable. Figure 2 shows substantial mean differences in temporal demand between Map Matching 1 and 2 and between the two Reading Comprehension tasks; Section 4.3.3 explicitly claims these pairs 'share demand level,' which the data do not support. Under absolute-agreement ICC, such column differences reduce ICC, and after per-speaker z-scoring the demand mismatch can drive the normalized ICC negative even when a real within-speaker load signal exists. A demand-matched reanalysis would settle this. This does not overturn the conditional verdict, but it identifies a concrete robustness check that should be run before the paper's headline negative result is treated as established. I also note the paper provides no code/data artifacts and no confidence intervals for the ICC changes, but those are secondary to the conceptual confound addressed here.","tokens_in":15981,"tokens_out":12911,"duration_ms":135630,"concrete_test":"Recompute Section 5.3 / Table 4 normalized ICC using only dyads where the two members of each retest pair have approximately matched self-reported NASA-TLX demand, e.g. |temporal_demand(task4) - temporal_demand(task9)| ≤ 2 and similarly for tasks 6/7, or residualize within-speaker z-scores on task-column demand ratings before computing ICC. If normalized acoustic ICC rises appreciably above 0 (or the raw-to-normalized drop shrinks), the Section 5.3 conclusion is an artifact of demand mismatch; if it remains near or below 0, the identity-artifact interpretation survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 5.3) interprets the drop from raw acoustic ICC 0.576 to normalized -0.112 as evidence that prosodic features 'measure who speaks rather than how speakers behave under load.' This interpretation requires the two test-retest pairs (map-matching tasks 4 and 9; reading-comprehension tasks 6 and 7) to be exchangeable with respect to the state being measured. The paper states they 'share task type and demand level' (Section 4.3.3), but Figure 2's task means contradict this: Map Matching 1 shows temporal demand around 15.8 while Map Matching 2 shows around 10.3; Reading Comp 1 and Reading Comp 2 also differ on this scale. Because ICC(2,1) is absolute agreement, a systematic demand difference between the two columns lowers ICC. After per-speaker z-scoring across all nine tasks, a common demand-driven shift between the two 'retest' tasks becomes a column-mean difference, pushing ICC negative even if acoustic features carry a genuine within-speaker load response. The observed normalized value (-0.112) is also very close to the expected correlation (-1/(N-1) ≈ -0.125) for two randomly chosen coordinates from an N=9 per-speaker z-score null, so the negative sign is not by itself evidence of an identity artifact. The analysis therefore needs demand-matched retest pairs before the collapse can be attributed to speaker identity rather than task mismatch.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a three-dimensional evaluation framework for multimodal conversational-state features: predictive accuracy (leave-one-dyad-out regression/classification), cross-task generalizability (leave-one-category-out), and test-retest reliability (ICC(2,1)) before and after within-speaker z-score normalization. Using the AVCAffe dataset (53 dyads, 9 collaborative tasks), it compares interaction, acoustic, and linguistic feature families for cognitive load (NASA-TLX mental and temporal demand) and conversational power. The main reported findings are: linguistic features achieve the highest within-dataset predictive accuracy for temporal demand but collapse in cross-task evaluation; raw acoustic ICC (mean 0.576) drops to -0.112 after speaker normalization, which the authors interpret as prosodic features measuring speaker identity rather than behavioral state; interaction features are the only family whose reliability is unaffected by normalization; and floor imbalance correlates with within-dyad load asymmetry, while power classification remains at chance.","tokens_in":16272,"tokens_out":6219,"duration_ms":64122,"significance":"If the acoustic-reliability result is correct, it would challenge a common practice in multimodal affective computing—reporting raw acoustic ICC as evidence of feature stability—and would be a useful empirical contribution. The three-dimensional evaluation itself, applied to a public dataset, is a commendable and reusable idea. The paper is careful in several respects: it provides bootstrap confidence intervals for the LODO prediction results, reports feature-level ICC breakdowns, and acknowledges limitations (fixed task order, ASR quality, single-task problem-solving cell). However, the central acoustic conclusion is currently under-supported because the two test-retest pairs appear not to be matched in task demand, and the normalized ICC result is numerically close to a null expectation. The reliability and LOCO results also lack uncertainty quantification, and the cognitive-load targets were selected post hoc based on agreement in the same dataset. These issues are fixable but must be addressed before the central claims can be accepted.","major_comments":[{"comment":"The central claim that prosodic features 'measure who speaks rather than how speakers behave under load' depends on the two test-retest pairs (tasks 4/9 and 6/7) being exchangeable in the state being measured. The text asserts they 'share task type and demand level,' but Figure 2 shows noticeable differences in mean NASA-TLX scores between the two map-matching instances and between the two reading-comprehension instances (e.g., temporal demand roughly 10 vs 15 for map matching). Since ICC(2,1) is absolute agreement, a systematic occasion-level demand difference lowers raw ICC; after within-speaker z-scoring across all nine tasks, such a common shift becomes a column-mean difference and can drive ICC negative even if acoustic features carry a genuine within-speaker load response. The observed normalized ICC (-0.112) is also close to the -1/(N-1) ≈ -0.125 expectation for two random coordin","section":"§5.3, §4.3.3"},{"comment":"ICC(2,1) is reported as a point value without confidence intervals. With 53 dyads and only two test-retest pairs, the 119% drop from 0.576 to -0.112 could be within sampling variability. Bootstrap or F-based intervals should accompany these estimates. The same applies to LOCO results in Table 3: the linguistic temporal-demand contrast (0.009 on Social vs 0.406 on Time-pressured) is load-bearing for the task-specificity claim but is reported without intervals, so the reader cannot assess the strength of that contrast.","section":"Table 4, §5.3"},{"comment":"The cognitive-load targets (mental and temporal demand) were 'selected for their moderate within-dyad agreement' after inspecting agreement in the same AVCAffe dataset, while frustration and physical demand were excluded because their agreement was low. This is post-hoc selection on the dependent variable and conditions every subsequent predictive and reliability estimate on the chosen targets. The selection should be reported as a design decision with a pre-registered criterion, or supplemented with a sensitivity analysis covering all six TLX dimensions.","section":"§3.2"},{"comment":"Floor imbalance is defined in Table 5 as a signed quantity, (t_A - t_B)/(t_A + t_B). RQ4 examines its correlation with |load_A - load_B|, an unsigned asymmetry. If the signed value is used, the correlation would depend on the arbitrary labeling of speakers A and B; if the absolute value was used, the text should state so explicitly. This is load-bearing for the floor-dominance contribution and requires clarification or a reanalysis with the absolute value.","section":"§5.4, Table 5"},{"comment":"LOCO holds out task categories but keeps the same 53 dyads in the training folds, so this is cross-task evaluation within the same speakers. It does not test generalization to new speakers. The claim that linguistic features 'are learning task-specific vocabulary patterns rather than a generalizable load signal' is plausible, but the framing 'cross-task generalizability' should be qualified as 'cross-task, within-speaker' to avoid overstating the scope of the result.","section":"§4.3.2, §5.2"}],"minor_comments":[{"comment":"There is a missing space in 'AVCAffedataset' (line 2 of the abstract).","section":"Abstract"},{"comment":"The phrase '119% reduction' is the relative reduction from 0.576 to -0.112; since the value crosses zero, the absolute change (-0.688) may be less misleading for readers.","section":"§5.3"},{"comment":"The feature 'estimated words/min' appears in both the acoustic and linguistic feature sets. In combined conditions, this creates a duplicate feature unless explicitly removed. The appendix or Section 4.1 should clarify how duplicates are handled.","section":"Tables 6 and 7"},{"comment":"The fixed task-order confound is acknowledged, but it also affects the within-speaker z-score normalization across all nine tasks: if task demand trends with session time, the normalization window absorbs an orderly time trend. This additional consequence of fixed task order should be noted in the limitations.","section":"§6.2"},{"comment":"The floor-imbalance result (r=0.315, p<0.001) is reported without stating how many feature-outcome correlations were tested or whether any multiple-comparison correction was applied. The p-value would survive Bonferroni for 14 interaction features, but the total number of tests and correction procedure should be reported for transparency.","section":"§5.4"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern lands: the acoustic ICC collapse is currently over-interpreted because the retest pairs may not be demand-matched. This is the key barrier to acceptance. If the authors can re-run the reliability analysis with demand-matched pairs or a task-level fixed effect, and report confidence intervals, the paper would be a solid empirical contribution. The post-hoc target selection also needs to be handled transparently, but it is not fatal. I do not see grounds for rejection, provided the central acoustic claim is re-supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper to know about: Chowdhury's three-dimensional evaluation of multimodal features—prediction, cross-task generalizability, test-retest reliability—on the AVCAffe dyadic remote-collaboration dataset. The genuinely useful parts are the framework and several empirical findings: linguistic features predict temporal demand well (CCC 0.545) but collapse under cross-task evaluation (0.009 on social tasks); power classification sits at chance; interaction features like floor imbalance keep modest but real reliability. Those are plausible and worth building on.\n\nThe soft spot is the central acoustic claim. Raw ICC for acoustic features is 0.576 and drops to -0.112 after within-speaker z-scoring. The paper interprets this as \"measure who speaks rather than how speakers behave under load.\" Two concerns. First, the test-retest pairs are not demand-matched: Figure 2 shows Map Matching 1 with mean temporal demand around 15.8 and Map Matching 2 around 10.3; Reading Comp 1 and 2 also differ. For ICC(2,1) absolute agreement, a systematic column-mean difference lowers the coefficient, and after z-scoring across all nine tasks that difference can push ICC negative even if acoustic features carry a genuine within-speaker load response. The observed -0.112 is also close to the expected correlation of two random coordinates from a 9-point z-scored vector, -1/8 ≈ -0.125, so the negative sign alone is not evidence of an identity artifact. Second, the z-score normalization assumes task-induced acoustic changes are common across speakers; if some speakers raise pitch under load and others lower it, the normalization erases a real state signal. The collapse needs to be re-run on demand-matched pairs, or with demand as a covariate, before the speaker-identity interpretation holds.\n\nOther issues are minor in comparison: ICC and LOCO estimates have no confidence intervals; the outcome targets were selected after checking within-dyad agreement in the same dataset; LOCO uses the same dyads in training and test folds (though the paper notes the single-task category); and the RQ4 correlation is uncorrected for multiple comparisons. No code or data artifacts are provided, so exact reproduction is impossible.\n\nBottom line: the paper offers a useful evaluation template and some solid negative results, but the headline acoustic-reliability finding is not yet established. It deserves a serious referee, but a revision that addresses the demand-matching confound is necessary. If the central claim survives after that, it's a meaningful result.\n\nRecommended: send to peer review with that demand-matching issue front and center.","headline":"The framework and some null results are worth a look, but the headline acoustic-reliability collapse is confounded by unmatched test-retest demand levels and needs a re-analysis.","tokens_in":16747,"tokens_out":6599,"would_cite":true,"duration_ms":53784,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that common acoustic prosodic features measure who is speaking rather than how speakers behave under load: after within-speaker normalization, their reliability collapses from moderate to null.","keywords":["cognitive load","conversational power","multimodal feature evaluation","intraclass correlation","speaker normalization","cross-task generalization","dyadic interaction","acoustic prosody"],"falsifier":"Run the same dyads through two matched task versions that differ only in experimentally imposed time pressure (order counterbalanced, per-speaker). If within-speaker z-scored pitch or spectral features classify the high-load condition above chance across speakers, or if the sign of pitch change under load is consistent within a speaker but opposite across speakers, then the ICC collapse reflects removal of speaker-specific state responses, not pure identity artifact.","tokens_in":15820,"feed_emoji":"🎙️","tokens_out":6856,"duration_ms":63734,"temperature":0.7,"pith_summary":"This paper argues that standard accuracy-based evaluation hides what multimodal features actually measure, and that reliability must be judged after removing each speaker's own baseline. Using nine collaborative tasks from 53 remote dyads, it finds that acoustic features' moderate test-retest reliability (mean ICC ≈ 0.576) collapses to ≈ −0.112 once each feature is z-scored within speaker, implying that common prosodic cues measure who is speaking rather than how a person behaves under load. Why it matters: prosodic features are widely used in cognitive-load and affective sensing, so if right, previously reported acoustic reliability in that literature is systematically inflated by speaker identity. The same analysis shows linguistic features predict temporal demand best within a task but fail on held-out task types, while interaction features—floor control, silences, interruptions—are the only family whose reliability survives normalization, and floor imbalance tracks which partner carries more mental load.","feed_headline":"Acoustic reliability collapses once speaker identity is controlled","feed_subtitle":"Raw ICC 0.576 drops to −0.112 after within-speaker z-scoring; only interaction cues survive.","key_machinery":"The load-bearing instrument is the comparison of intraclass correlation coefficients (ICC(2,1)) before and after within-speaker z-score normalization. For each speaker, each feature is standardized across that speaker's nine task observations, removing all between-speaker variance (vocal anatomy, habitual amplitude and style) while keeping task-to-task variation. Any raw ICC that collapses after this transformation is attributed to speaker identity; any ICC that survives—like the interaction features, which are dyadic ratios and differences and therefore cancel between-speaker levels by construction—is treated as genuinely reliable. The three-dimensional framework (LODO CCC for prediction, L","core_discovery":"The central discovery is a split among feature families once prediction, generalization, and speaker-normalized reliability are evaluated together. Raw acoustic features look moderately reliable (mean ICC = 0.576), but z-scoring each feature within each speaker across the nine task observations drops mean ICC to −0.112, with spectral features falling from near 0.95 to near zero or negative. The paper reads this as evidence that standard prosodic features encode stable vocal-tract and speaking-style traits rather than task-induced state. Linguistic features predict temporal demand best within-distribution (CCC = 0.545) but collapse to 0.009 on held-out social tasks, indicating task-specific v","pith_inferences":["If the identity-inflation result holds beyond Zoom (e.g., in-person or telephone corpora), a large share of single-task acoustic load findings may need re-analysis with within-speaker normalization to separate trait from state; the paper itself flags audio compression and platform AGC as potential contributors.","The near-zero normalized acoustic ICC is only interpretable as 'no state signal' if task-induced acoustic changes are consistent in direction across speakers. A mixed-effects model with speaker-specific random slopes for task load could test whether some speakers raise and others lower pitch under load; if so, normalization would erase a real but idiosyncratic signal.","The interaction-feature result suggests a cheap deployment path: voice activity detection alone can feed floor-imbalance monitoring to rebalance participation in remote meetings, without ASR or speaker identification.","Extending the three-dimensional protocol to multi-party groups may restore power signals that dyadic structure suppresses, since dominance asymmetries are typically more observable with more participants."],"forward_implications":["Reported acoustic ICC values in multimodal affective computing should be recomputed within speaker; a drop larger than 50% means the reliability estimate reflects speaker-intrinsic voice traits, not behavioral consistency.","Linguistic features should not be treated as portable load detectors across task contexts; their high within-corpus accuracy for temporal demand is tied to task vocabulary, as shown by the collapse to 0.009 on social tasks.","Interaction features such as floor imbalance, silence, and interruption asymmetry are the only family whose reliability estimate survives speaker normalization, making them the defensible foundation for measurement systems even with modest predictive accuracy.","Floor imbalance can serve as a real-time, ASR-free indicator of which speaker in a dyad carries higher cognitive load, since it predicts within-dyad load asymmetry (r = 0.315).","Conversational power at task level appears unmeasurable from aggregated multimodal features in dyadic collaboration; all feature conditions sit at the chance baseline."],"fun_headline_variants":["Speaker normalization destroys acoustic reliability","Interaction features are the only reliable conversational cues","Linguistic cues predict load but fail cross-task","Power role prediction remains at chance in dyads"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The collapse of acoustic ICC after normalization is only damning if genuine load-driven acoustic changes live entirely in within-speaker residuals and are directionally consistent across speakers; if some speakers raise pitch and others lower it under load, removing each speaker's mean and spread erases real state signal along with the identity artifact.","fun_headline_variants_meta":{"raw":{"variants":["Speaker normalization destroys acoustic reliability","Interaction features are the only reliable conversational cues","Linguistic cues predict load but fail cross-task","Power role prediction remains at chance in dyads"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000178,"raw_usage":{"total_tokens":1160,"prompt_tokens":795,"completion_tokens":365,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":310}},"tokens_in":539,"tokens_out":365,"duration_ms":4099,"temperature":1.0,"reasoning_tokens":310,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T17:53:51.184545+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same dyads through two matched task versions that differ only in experimentally imposed time pressure (order counterbalanced, per-speaker). If within-speaker z-scored pitch or spectral features classify the high-load condition above chance across speakers, or if the sign of pitch change under load is consistent within a speaker but opposite across speakers, then the ICC collapse reflects removal of speaker-specific state responses, not pure identity artifact.","supporting_citations":[],"review_version":1}