REVIEW 3 major objections 4 minor 79 references
CARE: A Multimodal Corpus for Studying Speech and Non-Verbal Communication Across Multiple Medical Conditions
T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read This paper introduces CARE v1.0, a 144-hour multimodal video corpus of 612 speakers across 12 medical conditions plus a control group, and argues it is the first public benchmark enabling cross-disease speech and non-verbal behaviour analys
desk verdict Useful multi-condition multimodal resource, but its disease labels are HEXI topic labels, not verified diagnoses, so the 'disease detection' framing outruns the data. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The corpus itself is the load-bearing object: 4,281 short video clips organized as GROUP/SUBGROUP/PROFILE/VIDEO, with each profile linked to demographic and clinical metadata and each video accompanied by turn-level acoustic features, frame-wise facial landmarks, action units, gaze vectors, head pose, upper-body keypoints, and transcript-derived linguistic descriptors. The design that makes it work is the pairing of automatically extracted structured metadata with duplicate-profile links, so the same person can appear as a patient and as a caregiver, enabling comorbidity and cross-role analyses. The extraction pipeline uses speaker diarization to isolate the participant's turns and a fixed-t
What would settle it
Take a random subset of profiles and compare their platform group labels against independent clinical records or structured diagnostic interviews; if a substantial share of labels are wrong, or if a matched-pair analysis shows that age, sex, and caregiving role alone can separate patients from controls at chance-adjusted levels, then classifiers trained on the corpus would be detecting demographics and role rather than disease.
Extended reading notes
Core claim
The central claim is that CARE v1.0 establishes a new kind of resource for health-focused multimodal speech research: a curated corpus of genuine patient and caregiver interviews combined with dense behavioral features and clinically relevant metadata. The authors show that their curation pipeline can take a public health-experience archive, select 12 conditions with known or expected speech and non-verbal impact, filter profiles for usable video, and enrich each profile with structured fields such as current medication, comorbidities, dominant emotions, and caregiving role. They further validate that a large language model can extract these metadata fields with an average BERTScore of 0.75
Load-bearing premise
The load-bearing premise is that the health-condition labels attached to each profile are correct diagnoses and that the control group—caregivers, bereaved relatives, and professionals—is a valid comparison population; if either fails, any disease-detection benchmark built on this corpus inherits the error.
Editorial extensions
If this is right
- A single benchmark now supports automatic disease and symptom detection across 12 conditions, removing the need to stitch together incompatible single-disease datasets.
- Researchers can model speech and non-verbal behaviour under emotionally charged health communication, using both narrative-derived emotion labels and face- and voice-based affective trajectories.
- The documented medication, comorbidity, and demographic metadata allow studies that control for confounds when comparing conditions or patients against controls.
- The corpus supports longitudinal and coping-process analyses through time-since-onset, symptom, and life-impact fields extracted from the narratives.
- The released feature extraction scripts and example evaluation workflow promote reproducible, fair comparisons across future studies.
Reading between the lines
- Because the control profiles are caregivers, bereaved relatives, and professionals rather than healthy age- and sex-matched speakers, a classifier that separates patients from controls may partly learn caregiving-related emotional and conversational patterns rather than disease-specific signals; explicit role and emotional-state covariates are needed in benchmark designs.
- The emotion metadata comes from written narratives, not from the videos' facial or vocal signals; comparing narrative-derived emotions with the provided facial and acoustic emotion trajectories could reveal systematic differences between reported and expressed affect.
- The corpus's uneven condition sizes make it a natural testbed for transfer learning: models trained on well-represented conditions such as Parkinson's or depression might improve detection in low-resource groups like cleft lip and palate if cross-disease speech and movement patterns exist.
- Independent clinical verification of a subset of the source profiles would materially strengthen the benchmark; without it, label-noise sensitivity analyses are a concrete and sensible next step.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces CARE v1.0, a curated multimodal dataset derived from the Health Experiences (HEXI) platform. It comprises 622 participant profiles (612 unique individuals) across 12 condition-related groups plus a control group, yielding 4,281 short video clips totaling approximately 143.5 hours. The dataset ships pre-computed audio, facial, gaze, body-movement, and linguistic descriptors, along with structured metadata (demographics, medications, symptoms, life impacts, and emotions) extracted by a GPT model. The authors validate this extraction on 34 manually annotated profiles, compare several LLMs, and provide feature-extraction scripts and an example evaluation workflow. The stated goal is to enable multimodal research on disease and symptom detection and on non-verbal communication in health contexts.
Significance. If the condition labels and control group were valid as framed, CARE would be a valuable first public multi-condition multimodal corpus, addressing a well-known gap in single-condition, audio-only datasets. The breadth of features, the diarization-based segmentation, the transparency of the duplication flag, and the release of processing scripts are concrete strengths. The paper also provides an honest technical validation of the LLM metadata. However, the scientific contribution for disease detection depends critically on the diagnostic validity of the group labels and the comparability of the control group, both of which are explicitly and inherently limited by the source archive. The value of the corpus for narrative and health-experience research is real, but the advertised disease-detection framing overreaches the evidence.
major comments (3)
- [Section 3, GROUP definition] The paper states that GROUP 'does not necessarily correspond to the clinical label of the speaker' and that the same participant can appear as a patient in one profile and as a caregiver in another. Since HEXI is an experience-sharing archive, the 12 condition labels are topic-level categories rather than clinically verified diagnoses. The abstract and contributions nevertheless characterize CARE as spanning '12 medical conditions' and enabling 'automatic disease and symptom detection.' This is an overclaim unless label validity is quantified. I recommend either verifying diagnoses on a held-out subset (e.g., against clinical records), or systematically rephrasing 'condition' as 'health-experience topic' and adding a prominent limitation in the abstract and usage notes.
- [Section 4, Figure 3] The 'control' group is composed of caregivers, bereaved relatives, advocates, and healthcare/research professionals, not matched healthy controls. Any patient-vs-control comparison risks capturing narrative perspective (first-person vs. third-person), emotional state, age/sex distribution, or recruitment context rather than the condition itself. The claim that this is 'a more contextually relevant reference population' does not address confounding. The authors should either rename this group (e.g., 'non-patient experience narrators'), restrict the benchmark claims to narrative/topic classification, or provide a covariate-matched analysis to demonstrate that the control group is usable for disease-detection benchmarks.
- [Section 5, Table 3] The LLM validation is based on only 34 profiles annotated by a single non-clinician annotator. Several fields central to the advertised 'comprehensive' metadata show poor agreement: 'All medications or treatments mentioned' has ACC 0.00 and JS at most 0.15 across all models; 'Dominant emotions' has ACC 0.29 (patient) / 0.40 (control) and JS 0.46/0.47; numerous symptom and psychosocial fields have ACC below 0.30. While BERTScore is higher and the authors are transparent about the metric limitations, the released dataset still includes these labels as structured metadata. I recommend flagging the least reliable fields as provisional, withholding them, or substantially expanding the validation set with multiple annotators and inter-annotator reliability metrics. As it stands, the 'structured metadata covering medication, life impacts, and emotions' claim is not fully supported.
minor comments (4)
- [Section 2.2] The exclusion criteria include 'AI-generated or synthetic looping footage' and 'recordings performed by actors.' Please describe how these were detected (manual inspection, automated screening, or platform labels).
- [Section 5] For 'years of care' and 'years since condition onset,' the paper says the MAE was 'equal or below one year' but does not report exact values or the number of numeric pairs used. Please provide these details.
- [Throughout] Minor typos: 'insufficent' in Section 1, 'Similary' in Section 3, and several stray spaces in references (e.g., 'Parkinson ’s'). A careful proofreading pass is needed.
- [Section 4, Figure 2] The middle panel uses a logarithmic scale for number of subjects; clarify in the caption that the scale is logarithmic, since the axis labels are raw counts.
Circularity Check
No circularity: the corpus construction and validation loops are independent of the claims being made.
full rationale
CARE is a data resource paper, not a derivational one. The central claim is that a curated multimodal corpus was assembled from the external HEXI archive. The construction pipeline (Section 2.2) selects and filters profiles by video usability, and the metadata extraction uses GPT-oss-120b. The model choice is independently validated in Section 5 against a manually annotated held-out subset of 34 profiles, and the reported metrics (ACC, Jaccard, BERTScore) are against human reference annotations, not against the downstream claims. No parameter is fitted to a subset of data and then reported as a prediction of that same subset; no equation is derived from its own outputs; no uniqueness theorem is imported from the authors' prior work to force a conclusion. The unverified HEXI group labels and the non-matched control cohort are validity and bias concerns external to circularity: the labels are inputs, not outputs of the paper's analysis. Therefore, no significant circularity is present and the score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Participant profiles from HEXI accurately represent clinically diagnosed patients and controls.
- domain assumption The 12 selected conditions have known or expected relevance to speech and non-verbal communication.
- domain assumption LLM-extracted metadata is sufficiently reliable for downstream clinical/computational use.
- domain assumption Automatic feature extraction tools (OpenFace, MediaPipe, DiariZen, WhisperX, etc.) produce accurate enough descriptors on heterogeneous videos.
Cite this review
Pith. "Pith review of CARE: A Multimodal Corpus for Studying Speech and Non-Verbal Communication Across Multiple Medical Conditions." pith.science (2026). https://pith.science/paper/6NTSS7OX
@misc{pith2026260725903,
author = {Pith},
title = {Pith review of: CARE: A Multimodal Corpus for Studying Speech and Non-Verbal Communication Across Multiple Medical Conditions},
year = {2026},
howpublished = {\url{https://pith.science/paper/6NTSS7OX}},
note = {Machine review of arXiv:2607.25903}
}
read the original abstract
Automatic analysis of multimodal speech has shown strong potential for computationally detecting and monitoring a wide range of neurological, psychiatric, and respiratory conditions. However, progress in this field is limited by existing publicly accessible datasets, which are often small in scale, focused on a single condition or disease, and primarily speech focused. Moreover, if key confounding variables such as education, medication use, comorbidities, or mood state are insufficiently documented, the reliability and interpretability of computational analyses are further compromised. To address these limitations, we introduce CARE v1.0, a curated multimodal English dataset of approximately 144 hours of short video interviews collected from 612 individuals across 12 medical conditions plus a control cohort. For each video, a comprehensive set of clinically relevant multimodal descriptors is provided, alongside structured metadata covering factors such as medication, life impacts, and expressed emotions. The corpus's breadth and heterogeneity support a wide range of applications, including automatic disease and symptom detection, multimodal modelling of speech and non-verbal behaviour under emotionally charged contexts, and studies of disease trajectories and coping processes.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
2020 , publisher=
Ziebland, Sue and Grob, Rachel and Schlesinger, Mark , journal=. 2020 , publisher=
2020
-
[2]
2004 , publisher=
Herxheimer, Andrew and Ziebland, Sue , journal=. 2004 , publisher=
2004
-
[3]
The Lancet , volume=
Database of patients' experiences (DIPEx): a multi-media approach to sharing experiences and information , author=. The Lancet , volume=. 2000 , publisher=
2000
-
[4]
and Trancoso, Isabel and Abad, Alberto , year=
Gimeno-Gómez, David and Botelho, Catarina and Martínez-Hinarejos, Carlos-D. and Trancoso, Isabel and Abad, Alberto , year=
-
[5]
Speech as a Biomarker for Disease Detection , year=
Botelho, Catarina and Abad, Alberto and Schultz, Tanja and Trancoso, Isabel , journal=. Speech as a Biomarker for Disease Detection , year=
-
[6]
, journal=
Gimeno-Gómez, David and Botelho, Catarina and Pompili, Anna and Abad, Alberto and Martínez-Hinarejos, Carlos-D. , journal=. Unveiling Interpretability in Self-Supervised Speech Representations for Parkinson's Diagnosis , year=
-
[7]
DNN embeddings: What to use in the automatic assessment of Parkinson’s disease in multi-lingual scenarios , author=
Interpretable speech features vs. DNN embeddings: What to use in the automatic assessment of Parkinson’s disease in multi-lingual scenarios , author=. Computers in Biology and Medicine , volume=. 2023 , publisher=
2023
-
[8]
doi:10.21437/Interspeech.2024-1255 , issn =
Catarina Botelho and John Mendonça and Anna Pompili and Tanja Schultz and Alberto Abad and Isabel Trancoso , year =. doi:10.21437/Interspeech.2024-1255 , issn =
Show all 79 references
-
[9]
Journal of Medical Internet Research , volume=
Automated Speech Markers of Alzheimer Dementia: Test of Cross-Linguistic Generalizability , author=. Journal of Medical Internet Research , volume=. 2025 , publisher=
2025
-
[10]
Journal of Alzheimer’s disease , volume=
Linguistic features identify Alzheimer’s disease in narrative speech , author=. Journal of Alzheimer’s disease , volume=. 2015 , publisher=
2015
-
[11]
Journal of Medical Internet Research , volume=
Applied machine learning techniques to diagnose voice-affecting conditions and disorders: systematic literature review , author=. Journal of Medical Internet Research , volume=. 2023 , publisher=
2023
-
[12]
Neurology , volume=
Speech and language delay are early manifestations of juvenile-onset Huntington disease , author=. Neurology , volume=. 2006 , publisher=
2006
-
[13]
Neurocomputing , volume=
Monitoring amyotrophic lateral sclerosis by biomechanical modeling of speech production , author=. Neurocomputing , volume=. 2015 , publisher=
2015
-
[14]
Speech communication , volume=
A review of depression and suicide risk assessment using speech analysis , author=. Speech communication , volume=. 2015 , publisher=
2015
-
[15]
Psychological medicine , volume=
Acoustic speech markers for schizophrenia-spectrum disorders: a diagnostic and symptom-recognition tool , author=. Psychological medicine , volume=. 2023 , publisher=
2023
-
[16]
Engineering in Medicine and Biology Society (EMBC), 2012 Annual International Conference of the IEEE , pages=
Speech analysis for mood state characterization in bipolar patients , author=. Engineering in Medicine and Biology Society (EMBC), 2012 Annual International Conference of the IEEE , pages=. 2012 , organization=
2012
-
[17]
ICASSP , pages=
Ecologically valid long-term mood monitoring of individuals with bipolar disorder using speech , author=. ICASSP , pages=. 2014 , organization=
2014
-
[18]
ICASSP , pages=
Speech as a biomarker for obstructive sleep apnea detection , author=. ICASSP , pages=. 2019 , organization=
2019
-
[19]
IEEE Journal of Selected Topics in Signal Processing , volume=
Modeling Obstructive Sleep Apnea Voices Using Deep Neural Network Embeddings and Domain-Adversarial Training , author=. IEEE Journal of Selected Topics in Signal Processing , volume=. 2019 , publisher=
2019
-
[20]
Frontiers in digital health , volume=
Predicting pulmonary function from the analysis of voice: a machine learning approach , author=. Frontiers in digital health , volume=. 2022 , publisher=
2022
-
[21]
and Schuller, Dagmar M
Schuller, Björn W. and Schuller, Dagmar M. and Qian, Kun and Liu, Juan and Zheng, Huaiyuan and Li, Xiao , TITLE=. Frontiers in Digital Health , VOLUME=. 2021 , DOI=
2021
-
[22]
2011 , booktitle =
The INTERSPEECH 2011 speaker state challenge , author =. 2011 , booktitle =. doi:10.21437/Interspeech.2011-801 , issn =
2011 doi
-
[23]
Computer Speech & Language , volume=
Medium-term speaker states—A review on intoxication, sleepiness and the first challenge , author=. Computer Speech & Language , volume=. 2014 , publisher=
2014
-
[24]
Schuller, Bj. The. Interspeech , year=
-
[25]
The INTERSPEECH 2017 computational paralinguistics challenge: Addressee, cold & snoring , author =
2017
-
[26]
Björn Schuller and Stefan Steidl and Anton Batliner and Simone Hantke and Florian Hönig and J. R. Orozco-Arroyave and Elmar Nöth and Yue Zhang and Felix Weninger , year =. doi:10.21437/Interspeech.2015-179 , issn =
2015 doi
-
[27]
Schuller and Anton Batliner and Christian Bergler and Florian B
Björn W. Schuller and Anton Batliner and Christian Bergler and Florian B. Pokorny and Jarek Krajewski and Margaret Cychosz and Ralf Vollmann and Sonja-Dana Roelen and Sebastian Schnieder and Elika Bergelson and Alejandrina Cristia and Amanda Seidl and Anne S. Warlaumont and Li...
2019 doi
-
[28]
Interspeech , pages=
The interspeech 2017 computational paralinguistics challenge: Addressee, cold & snoring , author=. Interspeech , pages=
2017
-
[29]
2025 , note =
Bensoussan, Yael and Sigaras, Alexandros and Rameau, Anais and Elemento, Olivier and Powell, Maria and Dorr, David and Payne, Philip and Ravitsky, Vardit and Bélisle-Pipon, Jean-Christophe and Bahr, Ruth and Watts, Stephanie and Bolser, Donald and Siu, Jennifer and Lerner-Elli...
2025 doi
-
[30]
Proceedings of the 6th International Workshop on Audio/Visual Emotion Challenge and Workshop , pages =
Valstar, Michel and Gratch, Jonathan and Schuller, Bj\". Proceedings of the 6th International Workshop on Audio/Visual Emotion Challenge and Workshop , pages =. 2016 , doi =
2016
-
[31]
IEEE Transactions on Audio, Speech, and Language Processing , publisher =
Automatic intonation recognition for the prosodic assessment of language-impaired children , author =. IEEE Transactions on Audio, Speech, and Language Processing , publisher =
-
[32]
and others , year = 2017, booktitle =
Ringeval, F. and others , year = 2017, booktitle =
2017
-
[33]
Proceedings of the 8th on Audio/Visual Emotion Challenge and Workshop , pages=
Ringeval, Fabien and Schuller, Bj. Proceedings of the 8th on Audio/Visual Emotion Challenge and Workshop , pages=
-
[34]
Proceedings of the 9th International on Audio/visual Emotion Challenge and Workshop , pages=
AVEC 2019 workshop and challenge: state-of-mind, detecting depression with AI, and cross-cultural affect recognition , author=. Proceedings of the 9th International on Audio/visual Emotion Challenge and Workshop , pages=
2019
-
[35]
npj Parkinson's Disease , volume=
Unveiling early signs of Parkinson’s disease via a longitudinal analysis of celebrity speech recordings , author=. npj Parkinson's Disease , volume=. 2024 , publisher=
2024
-
[36]
Saturnino Luz and Fasih Haider and Sofia de la Fuente and Davida Fromm and Brian MacWhinney , year =
-
[37]
An Overview of the ADReSS-M Signal Processing Grand Challenge on Multilingual Alzheimer's Dementia Recognition Through Spontaneous Speech , year=
Luz, Saturnino and Haider, Fasih and Fromm, Davida and Lazarou, Ioulietta and Kompatsiaris, Ioannis and MacWhinney, Brian , journal=. An Overview of the ADReSS-M Signal Processing Grand Challenge on Multilingual Alzheimer's Dementia Recognition Through Spontaneous Speech , year=
-
[38]
2024 , booktitle =
Saturnino Luz and Sofia. 2024 , booktitle =
2024
-
[39]
ICASSP , pages=
Early dementia detection using multiple spontaneous speech prompts: The process challenge , author=. ICASSP , pages=. 2025 , organization=
2025
-
[40]
Scientific Data , volume=
Mendes-Laureano, Jana. Scientific Data , volume=. 2024 , publisher=
2024
-
[41]
Tao, Fuxiang and Esposito, Anna and Vinciarelli, Alessandro , journal=
-
[42]
Haulcy, R’mani and Glass, James , booktitle=
-
[43]
Proceedings of the 29th Symposium on Operating Systems Principles , pages =
Kwon, Woosuk and Li, Zhuohan and Zhuang, Siyuan and Sheng, Ying and Zheng, Lianmin and Yu, Cody Hao and Gonzalez, Joseph and Zhang, Hao and Stoica, Ion , title =. Proceedings of the 29th Symposium on Operating Systems Principles , pages =. 2023 , doi =
2023
-
[44]
2025 , eprint=
gpt-oss-120b & gpt-oss-20b Model Card , author=. 2025 , eprint=
2025
-
[45]
arXiv preprint arXiv:2302.13971 , year=
Llama: Open and efficient foundation language models , author=. arXiv preprint arXiv:2302.13971 , year=
-
[46]
Qwen2.5: A Party of Foundation Models , url =
Qwen Team , month =. Qwen2.5: A Party of Foundation Models , url =
-
[47]
arXiv preprint arXiv:2407.10671 , year=
Qwen2 Technical Report , author=. arXiv preprint arXiv:2407.10671 , year=
-
[48]
2025 , eprint=
Qwen3 Technical Report , author=. 2025 , eprint=
2025
-
[49]
arXiv preprint arXiv:2501.15383 , year=
Qwen2.5-1M Technical Report , author=. arXiv preprint arXiv:2501.15383 , year=
-
[50]
1992 , publisher=
Hand and mind: What gestures reveal about thought , author=. 1992 , publisher=
1992
-
[51]
2004 , publisher=
Gesture: Visible action as utterance , author=. 2004 , publisher=
2004
-
[52]
2023 , volume=
Turaev, Sherzod and Al-Dabet, Saja and Babu, Aiswarya and Rustamov, Zahiriddin and Rustamov, Jaloliddin and Zaki, Nazar and Mohamad, Mohd Saberi and Loo, Chu Kiong , journal=. 2023 , volume=
2023
-
[53]
2027 , issn =
Computer Speech & Language , volume =. 2027 , issn =. doi:https://doi.org/10.1016/j.csl.2026.101990 , author =
2027
-
[54]
2014 , issn =
Speech Communication , volume =. 2014 , issn =. doi:https://doi.org/10.1016/j.specom.2013.09.008 , author =
2014 doi
-
[55]
Survey on Emotional Body Gesture Recognition , year=
Noroozi, Fatemeh and Corneanu, Ciprian Adrian and Kamińska, Dorota and Sapiński, Tomasz and Escalera, Sergio and Anbarjafari, Gholamreza , journal=. Survey on Emotional Body Gesture Recognition , year=
-
[56]
2025 , issn =
Computers in Biology and Medicine , volume =. 2025 , issn =. doi:https://doi.org/10.1016/j.compbiomed.2025.110896 , author =
2025
-
[57]
European Conference on Information Retrieval , pages=
Gimeno-G. European Conference on Information Retrieval , pages=. 2024 , organization=
2024
-
[58]
Computational and mathematical methods in medicine , volume=
Beltr. Computational and mathematical methods in medicine , volume=. 2018 , publisher=
2018
-
[59]
2013 , publisher=
Fiquer, Juliana Teixeira and Boggio, Paulo Sergio and Gorenstein, Clarice , journal=. 2013 , publisher=
2013
-
[60]
2024 , issn =
Journal of Psychiatric Research , volume =. 2024 , issn =. doi:https://doi.org/10.1016/j.jpsychires.2024.05.056 , author =
2024 doi
-
[61]
Cognition & emotion , volume=
An argument for basic emotions , author=. Cognition & emotion , volume=. 1992 , publisher=
1992
-
[62]
Weinberger and Yoav Artzi , booktitle=
Tianyi Zhang and Varsha Kishore and Felix Wu and Kilian Q. Weinberger and Yoav Artzi , booktitle=. 2020 , url=
2020
-
[63]
and Schuller, Björn W
Eyben, Florian and Scherer, Klaus R. and Schuller, Björn W. and Sundberg, Johan and André, Elisabeth and Busso, Carlos and Devillers, Laurence Y. and Epps, Julien and Laukka, Petri and Narayanan, Shrikanth S. and Truong, Khiet P. , journal=. 2016 , volume=. doi:10.1109/TAFFC.2...
2016
-
[64]
Vásquez-Correa, J. C. , title=. GitHub Repository , howpublished=. 2013 , publisher=
2013
-
[65]
Babu and C
A. Babu and C. Wang and A. Tjandra and K. Lakhotia and Q. Xu and N. Goyal and K. Singh and P. 2022 , booktitle=. doi:10.21437/Interspeech.2022-143 , issn=
2022 doi
-
[66]
Joel Shor and Subhashini Venugopalan , year =. Proc. Interspeech , pages =. doi:10.21437/Interspeech.2022-118 , issn =
2022 doi
-
[67]
IEEE Transactions on Pattern Analysis and Machine Intelligence , volume=
Dawn of the transformer era in speech emotion recognition: closing the valence gap , author=. IEEE Transactions on Pattern Analysis and Machine Intelligence , volume=. 2023 , publisher=
2023
-
[68]
2023 , booktitle=
Alexis Plaquet and Hervé Bredin , title=. 2023 , booktitle=
2023
-
[69]
Han, Jiangyu and Landini, Federico and Rohdin, Johan and Silnova, Anna and Diez, Mireia and Burget, Luk. Proc. ICASSP , year=
-
[70]
IEEE International Conference on Automatic Face and Gesture Recognition , year=
Baltru. IEEE International Conference on Automatic Face and Gesture Recognition , year=
-
[71]
2023 , url=
Zeng, Wenzheng and Xiao, Yang and Wei, Sicheng and Gan, Jinfang and Zhang, Xintao and Cao, Zhiguo and Fang, Zhiwen and Zhou, Joey Tianyi , booktitle=. 2023 , url=
2023
-
[72]
Soukupov. Proc. Computer Vision Winter Workshop (CVWW) , year=
-
[73]
and Sukthankar, Rahul and Sminchisescu, Cristian , booktitle=
Xu, Hongyi and Bazavan, Eduard Gabriel and Zanfir, Andrei and Freeman, William T. and Sukthankar, Rahul and Sminchisescu, Cristian , booktitle=. 2020 , volume=
2020
-
[74]
2021 , publisher=
Toisoul, Antoine and Kossaifi, Jean and Bulat, Adrian and Tzimiropoulos, Georgios and Pantic, Maja , journal=. 2021 , publisher=
2021
-
[75]
2021 , issn =
Biomedical Signal Processing and Control , volume =. 2021 , issn =. doi:https://doi.org/10.1016/j.bspc.2021.102612 , author =
2021
-
[76]
2023 , issn =
Computer Speech & Language , volume =. 2023 , issn =. doi:https://doi.org/10.1016/j.csl.2023.101519 , author =
2023
-
[77]
2008 , publisher=
Support vector machines , author=. 2008 , publisher=
2008
-
[78]
Max Bain and Jaesung Huh and Tengda Han and Andrew Zisserman , year =
-
[79]
2020 , url=
Song, Kaitao and Tan, Xu and Qin, Tao and Lu, Jianfeng and Liu, Tie-Yan , journal=. 2020 , url=
2020
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.