Pith. sign in

REVIEW 5 major objections 5 minor 34 references

This paper claims that common acoustic prosodic features measure who is speaking rather than how speakers behave under load: after within-speaker normalization, their reliability collapses from moderate to null.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 17:53 UTC pith:DGRB6HWX

load-bearing objection The framework and some null results are worth a look, but the headline acoustic-reliability collapse is confounded by unmatched test-retest demand levels and needs a re-analysis. the 5 major comments →

arxiv 2607.17452 v1 pith:DGRB6HWX submitted 2026-07-20 cs.CL

How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks

classification cs.CL
keywords cognitive loadconversational powermultimodal feature evaluationintraclass correlationspeaker normalizationcross-task generalizationdyadic interactionacoustic prosody
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that standard accuracy-based evaluation hides what multimodal features actually measure, and that reliability must be judged after removing each speaker's own baseline. Using nine collaborative tasks from 53 remote dyads, it finds that acoustic features' moderate test-retest reliability (mean ICC ≈ 0.576) collapses to ≈ −0.112 once each feature is z-scored within speaker, implying that common prosodic cues measure who is speaking rather than how a person behaves under load. Why it matters: prosodic features are widely used in cognitive-load and affective sensing, so if right, previously reported acoustic reliability in that literature is systematically inflated by speaker identity. The same analysis shows linguistic features predict temporal demand best within a task but fail on held-out task types, while interaction features—floor control, silences, interruptions—are the only family whose reliability survives normalization, and floor imbalance tracks which partner carries more mental load.

Core claim

The central discovery is a split among feature families once prediction, generalization, and speaker-normalized reliability are evaluated together. Raw acoustic features look moderately reliable (mean ICC = 0.576), but z-scoring each feature within each speaker across the nine task observations drops mean ICC to −0.112, with spectral features falling from near 0.95 to near zero or negative. The paper reads this as evidence that standard prosodic features encode stable vocal-tract and speaking-style traits rather than task-induced state. Linguistic features predict temporal demand best within-distribution (CCC = 0.545) but collapse to 0.009 on held-out social tasks, indicating task-specific v

What carries the argument

The load-bearing instrument is the comparison of intraclass correlation coefficients (ICC(2,1)) before and after within-speaker z-score normalization. For each speaker, each feature is standardized across that speaker's nine task observations, removing all between-speaker variance (vocal anatomy, habitual amplitude and style) while keeping task-to-task variation. Any raw ICC that collapses after this transformation is attributed to speaker identity; any ICC that survives—like the interaction features, which are dyadic ratios and differences and therefore cancel between-speaker levels by construction—is treated as genuinely reliable. The three-dimensional framework (LODO CCC for prediction, L

Load-bearing premise

The collapse of acoustic ICC after normalization is only damning if genuine load-driven acoustic changes live entirely in within-speaker residuals and are directionally consistent across speakers; if some speakers raise pitch and others lower it under load, removing each speaker's mean and spread erases real state signal along with the identity artifact.

What would settle it

Run the same dyads through two matched task versions that differ only in experimentally imposed time pressure (order counterbalanced, per-speaker). If within-speaker z-scored pitch or spectral features classify the high-load condition above chance across speakers, or if the sign of pitch change under load is consistent within a speaker but opposite across speakers, then the ICC collapse reflects removal of speaker-specific state responses, not pure identity artifact.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Reported acoustic ICC values in multimodal affective computing should be recomputed within speaker; a drop larger than 50% means the reliability estimate reflects speaker-intrinsic voice traits, not behavioral consistency.
  • Linguistic features should not be treated as portable load detectors across task contexts; their high within-corpus accuracy for temporal demand is tied to task vocabulary, as shown by the collapse to 0.009 on social tasks.
  • Interaction features such as floor imbalance, silence, and interruption asymmetry are the only family whose reliability estimate survives speaker normalization, making them the defensible foundation for measurement systems even with modest predictive accuracy.
  • Floor imbalance can serve as a real-time, ASR-free indicator of which speaker in a dyad carries higher cognitive load, since it predicts within-dyad load asymmetry (r = 0.315).
  • Conversational power at task level appears unmeasurable from aggregated multimodal features in dyadic collaboration; all feature conditions sit at the chance baseline.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the identity-inflation result holds beyond Zoom (e.g., in-person or telephone corpora), a large share of single-task acoustic load findings may need re-analysis with within-speaker normalization to separate trait from state; the paper itself flags audio compression and platform AGC as potential contributors.
  • The near-zero normalized acoustic ICC is only interpretable as 'no state signal' if task-induced acoustic changes are consistent in direction across speakers. A mixed-effects model with speaker-specific random slopes for task load could test whether some speakers raise and others lower pitch under load; if so, normalization would erase a real but idiosyncratic signal.
  • The interaction-feature result suggests a cheap deployment path: voice activity detection alone can feed floor-imbalance monitoring to rebalance participation in remote meetings, without ASR or speaker identification.
  • Extending the three-dimensional protocol to multi-party groups may restore power signals that dyadic structure suppresses, since dominance asymmetries are typically more observable with more participants.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper proposes a three-dimensional evaluation framework for multimodal conversational-state features: predictive accuracy (leave-one-dyad-out regression/classification), cross-task generalizability (leave-one-category-out), and test-retest reliability (ICC(2,1)) before and after within-speaker z-score normalization. Using the AVCAffe dataset (53 dyads, 9 collaborative tasks), it compares interaction, acoustic, and linguistic feature families for cognitive load (NASA-TLX mental and temporal demand) and conversational power. The main reported findings are: linguistic features achieve the highest within-dataset predictive accuracy for temporal demand but collapse in cross-task evaluation; raw acoustic ICC (mean 0.576) drops to -0.112 after speaker normalization, which the authors interpret as prosodic features measuring speaker identity rather than behavioral state; interaction features are the only family whose reliability is unaffected by normalization; and floor imbalance correlates with within-dyad load asymmetry, while power classification remains at chance.

Significance. If the acoustic-reliability result is correct, it would challenge a common practice in multimodal affective computing—reporting raw acoustic ICC as evidence of feature stability—and would be a useful empirical contribution. The three-dimensional evaluation itself, applied to a public dataset, is a commendable and reusable idea. The paper is careful in several respects: it provides bootstrap confidence intervals for the LODO prediction results, reports feature-level ICC breakdowns, and acknowledges limitations (fixed task order, ASR quality, single-task problem-solving cell). However, the central acoustic conclusion is currently under-supported because the two test-retest pairs appear not to be matched in task demand, and the normalized ICC result is numerically close to a null expectation. The reliability and LOCO results also lack uncertainty quantification, and the cognitive-load targets were selected post hoc based on agreement in the same dataset. These issues are fixable but must be addressed before the central claims can be accepted.

major comments (5)
  1. [§5.3, §4.3.3] The central claim that prosodic features 'measure who speaks rather than how speakers behave under load' depends on the two test-retest pairs (tasks 4/9 and 6/7) being exchangeable in the state being measured. The text asserts they 'share task type and demand level,' but Figure 2 shows noticeable differences in mean NASA-TLX scores between the two map-matching instances and between the two reading-comprehension instances (e.g., temporal demand roughly 10 vs 15 for map matching). Since ICC(2,1) is absolute agreement, a systematic occasion-level demand difference lowers raw ICC; after within-speaker z-scoring across all nine tasks, such a common shift becomes a column-mean difference and can drive ICC negative even if acoustic features carry a genuine within-speaker load response. The observed normalized ICC (-0.112) is also close to the -1/(N-1) ≈ -0.125 expectation for two random coordin
  2. [Table 4, §5.3] ICC(2,1) is reported as a point value without confidence intervals. With 53 dyads and only two test-retest pairs, the 119% drop from 0.576 to -0.112 could be within sampling variability. Bootstrap or F-based intervals should accompany these estimates. The same applies to LOCO results in Table 3: the linguistic temporal-demand contrast (0.009 on Social vs 0.406 on Time-pressured) is load-bearing for the task-specificity claim but is reported without intervals, so the reader cannot assess the strength of that contrast.
  3. [§3.2] The cognitive-load targets (mental and temporal demand) were 'selected for their moderate within-dyad agreement' after inspecting agreement in the same AVCAffe dataset, while frustration and physical demand were excluded because their agreement was low. This is post-hoc selection on the dependent variable and conditions every subsequent predictive and reliability estimate on the chosen targets. The selection should be reported as a design decision with a pre-registered criterion, or supplemented with a sensitivity analysis covering all six TLX dimensions.
  4. [§5.4, Table 5] Floor imbalance is defined in Table 5 as a signed quantity, (t_A - t_B)/(t_A + t_B). RQ4 examines its correlation with |load_A - load_B|, an unsigned asymmetry. If the signed value is used, the correlation would depend on the arbitrary labeling of speakers A and B; if the absolute value was used, the text should state so explicitly. This is load-bearing for the floor-dominance contribution and requires clarification or a reanalysis with the absolute value.
  5. [§4.3.2, §5.2] LOCO holds out task categories but keeps the same 53 dyads in the training folds, so this is cross-task evaluation within the same speakers. It does not test generalization to new speakers. The claim that linguistic features 'are learning task-specific vocabulary patterns rather than a generalizable load signal' is plausible, but the framing 'cross-task generalizability' should be qualified as 'cross-task, within-speaker' to avoid overstating the scope of the result.
minor comments (5)
  1. [Abstract] There is a missing space in 'AVCAffedataset' (line 2 of the abstract).
  2. [§5.3] The phrase '119% reduction' is the relative reduction from 0.576 to -0.112; since the value crosses zero, the absolute change (-0.688) may be less misleading for readers.
  3. [Tables 6 and 7] The feature 'estimated words/min' appears in both the acoustic and linguistic feature sets. In combined conditions, this creates a duplicate feature unless explicitly removed. The appendix or Section 4.1 should clarify how duplicates are handled.
  4. [§6.2] The fixed task-order confound is acknowledged, but it also affects the within-speaker z-score normalization across all nine tasks: if task demand trends with session time, the normalization window absorbs an orderly time trend. This additional consequence of fixed task order should be noted in the limitations.
  5. [§5.4] The floor-imbalance result (r=0.315, p<0.001) is reported without stating how many feature-outcome correlations were tested or whether any multiple-comparison correction was applied. The p-value would survive Bonferroni for 14 interaction features, but the total number of tests and correction procedure should be reported for transparency.

Circularity Check

0 steps flagged

No significant circularity; empirical decomposition, not a definitional or fitted-input reduction.

full rationale

This is an empirical reliability study rather than a derivation chain. The LODO prediction, LOCO generalizability, and ICC(2,1) reliability evaluations are three independent measurement operations; no parameter is fitted to the reliability target and then reported as a prediction. The central acoustic finding (raw ICC 0.576 to normalized -0.112) is not circular: within-speaker z-scoring removes between-speaker means and spreads, but whether the remaining within-speaker variation still shows test-retest agreement is an open empirical question, and the negative answer is a measured outcome. The attribution of the removed variance to speaker identity is an operational interpretation, not an equation-level identity. The interaction-feature claim that dyadic relative measures are 'structurally immune to speaker identity inflation' is a stated design property; Table 4 indicates interaction features were not normalized, so some text saying they are identical before/after is a minor reporting inconsistency, not a circular step. The skeptic's concern about unmatched demand levels in the two test-retest pairs is a construct-validity threat to the comparison, not a circularity: it questions comparability of retest conditions, not whether the result was presupposed. There is no load-bearing self-citation, no imported uniqueness theorem, and no fitted constant renamed as a prediction. Acknowledged limitations (fixed task order, ASR bias, single-task LOCO cell) are correctly stated and do not hide a circular derivation.

Axiom & Free-Parameter Ledger

4 free parameters · 6 axioms · 0 invented entities

The paper rests on measurement assumptions rather than invented entities: dyadic-mean NASA-TLX as ground truth, task-level aggregation, and the speaker-normalization decomposition. The latter is the load-bearing assumption; if it fails, the headline acoustic-reliability result weakens. The feature sets are standard and drawn from prior literature.

free parameters (4)
  • LOCO task-category grouping = Social (T1-2); Cognitive (T3,6,7); Problem-solving (T5); Time-pressured (T4,8,9)
    Hand-chosen grouping in Section 4.3.2 defines which held-out condition counts as cross-task; the linguistic collapse to 0.009 on Social tasks may change under a different grouping.
  • Within-speaker z-score normalization window = Each speaker's 9 task observations
    Central reliability result depends on this exact normalization; alternative ways of controlling speaker identity (e.g., mixed-effects partialling) could yield different magnitudes.
  • NASA-TLX target selection criterion = Mental and temporal demand selected; frustration (r=0.113) and physical demand (r=0.071) excluded; effort omitted as red
    Targets were chosen after inspecting within-dyad agreement in the same dataset (Section 3.2), a data-dependent selection that may inflate apparent predictive performance.
  • Interaction feature aggregation window = 30 seconds
    Interaction features are aggregated from 30-second windows to task level (Section 4.1); the window size is a free choice that affects floor imbalance, silence, and turn-taking features.
axioms (6)
  • domain assumption NASA-TLX dyadic mean is a valid target for cognitive load
    Section 3.3.1 uses the dyadic mean of mental/temporal demand ratings as the regression target; if the mean masks individual load differences or is dominated by one speaker, prediction results are affected.
  • domain assumption Task-level feature aggregation is sufficient
    Section 4.2 deliberately avoids window-level modeling; if cognitive load varies within a task and is averaged away, predictive and reliability estimates may be attenuated.
  • ad hoc to paper Within-speaker z-score normalization removes speaker identity while preserving behavioral state signal
    Section 4.3.3 defines normalization across each speaker's nine tasks; the central acoustic-reliability conclusion depends on this decomposition being valid.
  • domain assumption Distil-Whisper ASR transcripts are adequate for linguistic features
    Section 3.4 acknowledges ASR quality varies across the 18 countries of origin; linguistic feature predictions and reliability may inherit transcription errors.
  • standard math ICC(2,1) is appropriate with two repeated task instances per speaker/dyad
    Section 4.3.3 uses ICC(2,1) with two test-retest pairs; ICC estimates from two observations per target can be unstable and have wide confidence intervals.
  • domain assumption Random Forest on 53 dyads with 10-44 features yields stable feature-importance estimates
    Small n relative to feature count can lead to overfitting; bootstrap CIs are used, but no formal significance test is available for LODO-fold comparisons (Section 5.1).

pith-pipeline@v1.3.0-alltime-deepseek · 15660 in / 13462 out tokens · 130424 ms · 2026-08-01T17:53:51.184545+00:00 · methodology

0 comments
read the original abstract

Measuring conversational states such as cognitive load and conversational power from multimodal behavior requires characteristic features that are not only predictive but also reliable across task contexts. We present a three-dimensional evaluation framework assessing predictive accuracy, cross-task generalizability, and test-retest reliability, applied to interactional, acoustic, and linguistic features extracted from dyadic conversations during collaborative tasks performed over a video-conferencing platform (AVCAffe dataset; 53 dyads, 9 tasks). Our results show that no single feature family dominates all three dimensions. Linguistic features show the highest predictive accuracy for cognitive load but collapse under cross-task evaluation, revealing sensitivity to task-specific vocabulary. Additionally, acoustic reliability, often reported as evidence of feature stability, degrades once speaker identity is controlled, confirming that standard prosodic features measure vocal characteristics rather than conversational state. Interaction features provide the only genuinely reliable signal, unchanged after speaker normalization. Interestingly, classifying power role remained near chance baseline across all conditions, indicating limitations of task-level aggregated behavior for predicting power role in conversation. Our findings reveal three insights: (1) linguistic features predict best but generalize poorly across task contexts; (2) acoustic reliability collapses to near-zero once speaker identity is controlled, challenging standard evaluation practice; and (3) interaction features provide the only genuinely reliable signal, with floor dominance predicting within-dyad cognitive load asymmetry. These results argue for speaker normalization and multi-dimensional evaluation as prerequisites for context-aware, robust multimodal feature selection in conversational systems.

Figures

Figures reproduced from arXiv: 2607.17452 by Tahiya Chowdhury.

Figure 1
Figure 1. Figure 1: Our three-dimensional evaluation framework. Each [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: NASA-TLX cognitive load scores per task (mean ± std, 𝑛 = 53 dyads). Bar colors indicate task cate￾gory. Social tasks (Tasks 1–2, green) elicit the lowest load on both dimensions, explaining why the Social category is the most challenging held-out condition in cross-task evaluation. Map Matching 1 (Task 4) and Lost at Sea (Task 5) show the highest temporal demand. The gray dashed line marks the scale midpoi… view at source ↗
Figure 4
Figure 4. Figure 4: Top three features per family for each prediction [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Cross-task generalizability results. Bars show CCC when the column category is held out for testing. Dashed lines in [PITH_FULL_IMAGE:figures/full_fig_p006_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Test-retest reliability. Filled circles ( [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

34 extracted references · 7 canonical work pages

  1. [1]

    Jennifer Abel and Molly Babel. 2017. Cognitive Load Reduces Perceived Linguistic Convergence between Dyads.Language and Speech60, 3 (2017), 479–502. doi:10. 1177/0023830916665652

  2. [2]

    Boyer, P.V

    S. Boyer, P.V. Paubel, R. Ruiz, and R. El Yagoubi. 2018. Human Voice as a Measure of Mental Load Level.Journal of Speech, Language, and Hearing Research61, 6 (2018), 1403–1415. doi:10.1044/2018_JSLHR-S-18-0066

  3. [3]

    Cristian Danescu-Niculescu-Mizil, Lillian Lee, Bo Pang, and Jon Kleinberg. 2012. Echoes of power: language effects and power differences in social interaction. InProceedings of the 21st International Conference on World Wide Web(Lyon, France)(WWW ’12). Association for Computing Machinery, New York, NY, USA, 699–708. doi:10.1145/2187836.2187931

  4. [4]

    Ingy Farouk Emara and Nabil Hamdy Shaker. 2024. The impact of non- native English speakers’ phonological and prosodic features on automatic speech recognition accuracy.Speech Commun.157, C (Feb. 2024), 12 pages. doi:10.1016/j.specom.2024.103038

  5. [5]

    Riccardo Fusaroli, Bahador Bahrami, Karsten Olsen, Andreas Roepstorff, Geraint Rees, Chris Frith, and Kristian Tylén. 2012. Coming to Terms: Quantifying the Benefits of Linguistic Coordination.Psychological Science23, 8 (2012), 931–939. doi:10.1177/0956797612436816

  6. [6]

    Sanchit Gandhi, Patrick von Platen, and Alexander M. Rush. 2023. Distil- Whisper: Robust Knowledge Distillation via Large-Scale Pseudo Labelling. arXiv:2311.00430 [cs.CL]

  7. [7]

    Agustín Gravano and Julia Hirschberg. 2011. Turn-Taking Cues in Task-Oriented Dialogue.Computer Speech & Language25, 3 (2011), 601–634. doi:10.1016/j.csl. 2010.10.003

  8. [8]

    Tamil Selvan Gunasekaran, Kartik Gupta, Yun Suen Pai, Hao Bai, et al . 2025. CoAffinity: A Multimodal Dataset for Cognitive Load and Affect Assessment in Remote Collaboration.IEEE Transactions on Affective Computing16, 03 (2025). doi:10.1109/TAFFC.2025.3562559

  9. [9]

    Sandra G. Hart. 2006. NASA-Task Load Index (NASA-TLX): 20 years later. In Proceedings of the Human Factors and Ergonomics Society Annual Meeting, Vol. 50. 904–908. doi:10.1177/154193120605000909

  10. [10]

    Hart and Lowell E

    Sandra G. Hart and Lowell E. Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. InAdvances in Psychology. Vol. 52. Elsevier, 139–183. doi:10.1016/S0166-4115(08)62386-9

  11. [11]

    Mattias Heldner and Jens Edlund. 2010. Pauses, Gaps and Overlaps in Conversa- tions.Journal of Phonetics38, 4 (2010), 555–568. doi:10.1016/j.wocn.2010.08.002

  12. [12]

    Maximilian Ilse, Jakub Tomczak, and Max Welling. 2018. Attention-based Deep Multiple Instance Learning. InProceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 80), Jennifer Dy and Andreas Krause (Eds.). PMLR, 2127–2136. https://proceedings.mlr.press/ v80/ilse18a.html

  13. [13]

    Anne Johnstone, Umesh Berry, Tina Nguyen, and Alan Asper. 1995. There Was a Long Pause: Influencing Turn-Taking Behaviour in Human-Human and Human- Computer Spoken Dialogues.International Journal of Human-Computer Studies 42, 4 (1995), 383–411. doi:10.1006/ijhc.1995.1018

  14. [14]

    Asif Khawaja, Fang Chen, and Nadine Marcus

    M. Asif Khawaja, Fang Chen, and Nadine Marcus. 2012. Analysis of Collaborative Communication for Linguistic Cues of Cognitive Load.Human Factors54, 4 (2012), 518–529. doi:10.1177/0018720811431258

  15. [15]

    Rickford, Dan Jurafsky, and Sharad Goel

    Allison Koenecke, Andrew Nam, Emily Lake, Joe Nudell, Minnie Quartey, Zion Mengesha, Courtney Toups, John R. Rickford, Dan Jurafsky, and Sharad Goel. 2020. Racial Disparities in Automated Speech Recognition.Proceedings of the National Academy of Sciences117, 14 (2020), 7684–7689. doi:10.1073/pnas.1915768117

  16. [16]

    Koo and Mae Y

    Terry K. Koo and Mae Y. Li. 2016. A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research.Journal of Chiropractic Medicine15, 2 (2016), 155–163. doi:10.1016/j.jcm.2017.10.001

  17. [17]

    Levinson

    Stephen C. Levinson. 2016. Turn-Taking in Human Communication — Origins and Implications for Language Processing.Trends in Cognitive Sciences20, 1 (2016), 6–14. doi:10.1016/j.tics.2015.10.010

  18. [18]

    Lawrence I-Kuei Lin. 1989. A Concordance Correlation Coefficient to Evaluate Reproducibility.Biometrics45, 1 (1989), 255–268. doi:10.2307/2532051

  19. [19]

    Marianne Schmid Mast. 2002. Dominance as Expressed and Inferred Through Speaking Time: A Meta-Analysis.Human Communication Research28, 3 (2002), 420–450. doi:10.1111/j.1468-2958.2002.tb00814.x

  20. [20]

    Ngueajio and Gloria Washington

    Mikel K. Ngueajio and Gloria Washington. 2022. Hey ASR System! Why Aren’t You More Inclusive? Automatic Speech Recognition Systems’ Bias and Proposed Bias Mitigation Techniques. A Literature Review. InHCI International 2022 – Late Breaking Papers: Interacting with EXtended Reality and Artificial Intelligence: 24th International Conference on Human-Compute...

  21. [21]

    Schegloff, and Gail Jefferson

    Harvey Sacks, Emanuel A. Schegloff, and Gail Jefferson. 1974. A Simplest Sys- tematics for the Organization of Turn-Taking for Conversation.Language50, 4 (1974), 696–735. doi:10.2307/412243

  22. [22]

    Pritam Sarkar, Aaron Posen, and Ali Etemad. 2023. AVCAffe: A Large Scale Audio- Visual Dataset of Cognitive Load and Affect for Remote Work. InProceedings of the AAAI Conference on Artificial Intelligence (AAAI’23/IAAI’23/EAAI’23, Vol. 37). AAAI Press, Article 9, 10 pages. doi:10.1609/aaai.v37i1.25078

  23. [23]

    Björn Schuller, Stefan Steidl, Anton Batliner, Felix Burkhardt, Laurence Devillers, Christian Müller, and Shrikanth Narayanan. 2013. The INTERSPEECH 2013 Com- putational Paralinguistics Challenge: Social Signals, Conflict, Emotion, Autism. InProceedings of Interspeech. 148–152. doi:10.21437/Interspeech.2013-56

  24. [24]

    Björn Schuller, Stefan Steidl, Anton Batliner, Julien Epps, et al. 2014. The Inter- speech 2014 Computational Paralinguistics Challenge: Cognitive & Physical Load, Multitasking. InProceedings of Interspeech. 427–431. doi:10.21437/Interspeech. 2014-130

  25. [25]

    Ravid Shwartz-Ziv and Amitai Armon. 2022. Tabular Data: Deep Learning Is Not All You Need.Information Fusion81 (2022), 84–90. doi:10.1016/j.inffus.2021.11.011

  26. [26]

    Taptiklis, M

    N. Taptiklis, M. Su, J.H. Barnett, and C. Skirrow. 2023. Prediction of Mental Effort Derived from an Automated Vocal Biomarker Using Machine Learning in a Large-Scale Remote Sample.Frontiers in Artificial Intelligence6 (2023). doi:10.3389/frai.2023.1171652

  27. [27]

    Tausczik and James W

    Yla R. Tausczik and James W. Pennebaker. 2010. The Psychological Meaning of Words: LIWC and Computerized Text Analysis Methods.Journal of Language and Social Psychology29, 1 (2010), 24–54. doi:10.1177/0261927X09351676

  28. [28]

    Alessandro Vinciarelli, Maja Pantic, and Hervé Bourlard. 2010. Social Signal Processing: Survey of an Emerging Field.Signal Processing90, 5 (2010), 2130–2150. doi:10.1016/j.sigpro.2009.11.014

  29. [29]

    Maria Vukovic, Melissa Stolar, and Margaret Lech. 2021. Cognitive Load Estima- tion from Speech Commands to Simulated Aircraft.IEEE/ACM Transactions on Audio, Speech, and Language Processing29 (2021), 1011–1022. doi:10.1109/TASLP. 2021.3057492

  30. [30]

    H. Xu, L. Wang, J. Zou, J. Zhang, R. Li, et al. 2025. Recognising and Explaining Mental Workload Using Low-Interference Method by Fusing Speech, ECG and Eye Tracking Signals during Simulated Flight.Ergonomics(2025). doi:10.1080/ 00140139.2025.2511877

  31. [31]

    J. Yang, H. Yang, Z. Wu, and X. Wu. 2023. Cognitive Load Assessment of Air Traffic Controller Based on SCNN-TransE Network Using Speech Data.Aerospace 10, 7 (2023). doi:10.3390/aerospace10070584

  32. [32]

    Bo Yin, Fang Chen, and N. Ruiz. 2008. Speech-Based Cognitive Load Monitoring System. InIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2041–2044. doi:10.1109/ICASSP.2008.4518756 ICMI ’26, October 05–09, 2026, Napoli, Italy Chowdhury

  33. [33]

    Ruiz, Fang Chen, and M.A

    Bo Yin, N. Ruiz, Fang Chen, and M.A. Khawaja. 2007. Automatic Cognitive Load Detection from Speech Features. InProceedings of the 19th Australasian Conference on Computer-Human Interaction (OZCHI). Association for Computing Machinery, New York, NY, USA. doi:10.1145/1324892.1324946

  34. [34]

    Jianlong Zhou, Kun Yu, Fang Chen, Yang Wang, and Syed Zahid Arshad. 2018. Multimodal Behavioral and Physiological Signals as Indicators of Cognitive Load. InThe Handbook of Multimodal-Multisensor Interfaces. Vol. 2. ACM. doi:10.1145/ 3107990.3108002 Appendix: Feature Definitions Tables 5–7 provide definitions for all 44 unique features used in this study,...