REVIEW 4 major objections 6 minor 93 references
COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention
T0 review · 4 major / 6 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read This paper argues that perceived social intentions are inherently multiple and perspective-dependent, and presents COSI-Lab, a dataset that aligns participants' long-term goals with second-level, multi-observer 'apparent intention' annotati
desk verdict A genuinely useful dataset resource for social-intention research, but its core AII annotations are unvalidated, so the headline coupling claim is a promise rather than a demonstrated result. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Apparent Intent Inference (AII) problem: the task of inferring, from an ex-situ third-person perspective, what intention a person appears to have at a given moment, independent of whether that intention is later realized. The annotation protocol is the load-bearing instrument: it structures narratives through cues, situation characteristics, situation classes, and social scripts, and adds annotator-level trait data so that multiplicity of interpretations can be studied as a function of the perceiver rather than as noise.
What would settle it
Show that annotators' perceived intentions do not track any behavioral or self-report signal, for instance that participants' post-session goal-attainment ratings share no systematic association with the apparent intentions narrated by observers, or that two annotators' narratives for the same clip are no more similar to each other than to a shuffled baseline. More directly, if participants had reported their immediate proximal intentions during the event and observer narratives matched those reports at chance, the resource's validity as intention-perception data would collapse.
Extended reading notes
Core claim
The paper argues that in weakly scripted social settings, a person's apparent intention is not a single ground-truth label but a space of plausible interpretations that different perceivers construct from observable cues, situation assumptions, and their own interpretative tendencies. COSI-Lab operationalizes this by having crowd-sourced annotators watch 30-second overhead video clips of real mingling and write free-form intention narratives with timestamps, confidence, intensity, and counterfactual alternatives, while also capturing annotator traits and demographics. The resulting dataset is claimed to be the first to align these multi-perspective apparent-intention annotations with the obs
Load-bearing premise
The claim depends on ex-situ online annotators' free-form narratives being a valid window into genuine variation in intention perception, rather than arbitrary responses produced by poorly motivated crowd workers.
Editorial extensions
If this is right
- Researchers can study the relationship between self-reported long-term goals and apparent proximal intentions on the same individuals at the same event for the first time.
- Intention-aware systems can be trained or evaluated to output multiple plausible intention hypotheses with explanations, rather than being forced to commit to a single label.
- The dataset provides benchmark tasks for apparent intent inference and conversation group detection, including a human-LLM comparison that reveals systematic differences in how people and models describe intentions.
- Multi-perspective annotations coupled with annotator traits allow investigation of how demographics and reflective functioning shape intention perception, supporting perspectivist modeling.
- The weakly scripted, ecologically valid setting with real professional consequences makes findings more likely to transfer to in-the-wild social interactions than strongly scripted datasets.
Reading between the lines
- If the perspective-driven framing is correct, agreement between annotators should not be the primary quality metric; instead, downstream systems could be evaluated on the diversity and plausibility of the interpretations they generate, a departure from typical label-agreement benchmarks.
- The dataset could be extended with participants' own in-the-moment proximal intention self-reports, for example via brief experience sampling between sessions, to directly validate whether ex-situ observer narratives track the observed person's actual immediate intentions.
- The goal-to-perceived-intention link opens a route to study unrealized intentions: cases where observers perceive a goal-directed intention that the participant later reports not achieving, a key gap the paper identifies in existing intention research.
- Because annotator traits and demographics are collected, the data could support deliberately resampling perspectives to make intention-inference systems aware of who is perceiving, potentially reducing demographic bias in social AI.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. COSI-Lab is a multimodal dataset of a two-session conference mingling event involving 32 academics, with overhead video, close-talk and wearable audio, IMU, UWB, surveys, and annotations. The paper's main contribution is an annotation protocol that elicits multi-perspective 'apparent intention' narratives from crowd annotators alongside self-reported participant goals, and the claim that this is the first resource coupling long-term goals with short-term third-party perceived intentions in a weakly scripted social setting. The paper also presents descriptive analyses of conversation groups, topics, and goals, plus benchmark demonstrations for AII and conversation-group detection.
Significance. If the AII annotations are meaningful, the dataset is a potentially valuable resource for perspectivist, explainable intention-inference research: it includes synchronized multimodal data, privacy-preserving processing, an annotation protocol with theory-grounded components (cues, assumptions, scripts), and an honest LLM-as-judge evaluation that reports a 45.3% discrimination accuracy. The detailed data collection and processing description, camera-wise 5-fold group detection evaluation, and released code for reproducibility are strengths. However, the core value rests on the validity of the AII narratives, and that validity is not yet demonstrated.
major comments (4)
- [§5.1, §7.1, §8 (Cue Grounding)] The load-bearing claim—that COSI-Lab couples self-reported long-term goals with multi-perspective apparent-intention annotations—is not supported by any reliability or validity evidence for the AII annotations. There is no inter-annotator agreement metric, no verification that cues cited in the narratives are present in the video/audio, and no analysis relating the AII narratives to the observed participants' own self-reported goals. Section 8 states that cue grounding 'was not assessed' and that goals and AII 'bare some relationship' without testing it. This makes the central dataset contribution, and desiderata D1–D4 built on it, currently an unverified assertion. At minimum, the paper should report inter-annotator agreement (e.g., Krippendorff's alpha on intention categories or annotation components), a cue-presence check on a sample, and an exploratory comparison between AII narrativ
- [§6, Appendix E.2/E.3] The quantitative descriptive claims in Figure 4 (topic composition and topic-redirection uptake) are derived entirely from LLM-assisted coding with no human validation. The appendix provides detailed prompts, but no inter-coder reliability against human coders, and no two-human agreement baseline. As reported, these statistics are trustworthy only insofar as the Qwen3-14B labels are reliable, which is not established. The authors should validate a sample of the LLM coding against human coders and report agreement; otherwise these results should be presented as illustrative rather than quantitative findings.
- [§7.1] The 'intention estimation' benchmark is currently descriptive, not predictive. It compares human and Gemma4-generated narratives via semantic similarity and an LLM-as-judge accuracy of 45.3%, but no concrete task formulation, evaluation metric, or baseline protocol is given that future models could be measured against. The 45.3% figure is reported without variance or confidence intervals and without a significance test, so its interpretation is unclear. To claim a benchmark task, the paper should define an evaluation setup (e.g., predicting annotation components, matching a target perspective, or ranking narrative plausibility) and provide a reproducible metric.
- [§1, Table 1, §9] The claim of satisfying desideratum D4 ('hierarchical structure of human intention') and the conclusion that COSI-Lab 'for the first time' enables study of 'subgoals or apparent proximal intentions' go beyond what is demonstrated. The dataset contains self-reported long-term goals and AII for short clips, but no annotation explicitly linking proximal intentions to subgoals or to distal goal hierarchies. The checkmark for D4 in Table 1 and the conclusion wording should be softened or accompanied by an explicit description of which hierarchy-related variables are present in the released data.
minor comments (6)
- [Abstract] 'a explainable' should be 'an explainable'.
- [§5.1] 'the begging and end of the intention' should be 'the beginning and end'.
- [§8] 'they bare some relationship' should be 'they bear some relationship'.
- [Appendix C.2.1] 'felxible inputs' should be 'flexible inputs'.
- [Appendix B] 'the michrophones record' should be 'the microphones record'.
- [§7.1 / Figure 5] The semantic-similarity plots are described qualitatively as showing 'clear clustering separation'; a quantitative cluster-separation or distance statistic would make the claim more precise.
Circularity Check
No significant circularity: COSI-Lab is a resource paper whose benchmarks and analyses are self-contained or explicitly deferred.
full rationale
The paper's central contribution is the COSI-Lab dataset itself—a multimodal recording of a weakly scripted workshop with self-reported goals, multi-perspective AII annotations, and benchmark tasks. There is no derivation chain in which a predicted quantity is constructed from the same quantity it claims to predict, and no fitted parameter is renamed as a prediction. The AII annotations are elicited from crowd annotators and then compared with LLM outputs; the LLM-as-judge result (45.3%) is reported as a failure to distinguish human from model narratives, not as evidence for model superiority. The conversation-group detection benchmark trains standard models (DANTE, LSTM) on position/orientation features and evaluates against human group annotations; this is a conventional supervised evaluation, not circular. The descriptive topic-redirection analyses do rely on LLM-generated labels without human validation, which is a validity concern, but it is not circularity: the LLM is not being used to 'predict' a label that was itself derived from the LLM in a way that defines the result. The paper also openly concedes limitations that would be relevant to correctness or validity but not to circularity: cue grounding is not assessed, and the relationship between self-reported goals and AII is asserted rather than tested. Self-citations to ConfLab [46] and related prior work place the dataset in a lineage but do not function as an unverified load-bearing premise for the paper's conclusions. Overall, no step reduces to its own inputs by construction, so the circularity score is 0.
Assumptions & free parameters
free parameters (4)
- Assumed body height H =
1.7 m
- Keypoint height ratios α_i =
from COCO model (not given)
- LLM coding window parameters =
300-turn window, 250-turn stride (topic distribution); 300/220/80 for redirection screening
- Frame strides for baselines =
20 (LSTM), 300 (DANTE)
assumptions (6)
- domain assumption The workshop setting is ecologically valid for studying social intentions despite being recorded and weakly scripted.
- domain assumption Head orientation is a valid indicator of conversational attention and group membership (Kendon's F-formation).
- domain assumption 30-second clips are sufficient for a third-party observer to perceive proximal intentions.
- ad hoc to paper LLM-generated topic/redirection labels are acceptably reliable for descriptive statistics without human-validation metrics.
- domain assumption If multiple plausible intention narratives are socially meaningful, evaluating them requires diversity/grounding/plausibility measures.
- standard math Hyperbolic Tangent Similarity (HTS) is a valid semantic similarity metric.
Cite this review
Pith. "Pith review of COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention." pith.science (2026). https://pith.science/paper/C22Q3NW2
@misc{pith2026260728649,
author = {Pith},
title = {Pith review of: COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention},
year = {2026},
howpublished = {\url{https://pith.science/paper/C22Q3NW2}},
note = {Machine review of arXiv:2607.28649}
}
read the original abstract
COSI-Lab presents a multimodal, multi-sensor dataset of an interdisciplinary scientific workshop containing 32 academics at an international conference. It captures ecologically valid social interactions in a weakly scripted setting consisting of two 30-minute mingling sessions with real professional and social consequences for the participants involved. We argue that future intelligent systems could be better equipped to handle subjective perceptions by modeling their multiplicity not as label noise but as a explainable perspective-driven reasoning process. We focus on the Apparent Intent Inference (AII) problem as determined by ex-situ observers and conceptualize intentions to be independent of manifest future outcomes. We contribute 1. a novel annotation process for AII that accounts for a perceiver's own interpretative tendencies, 2. quantitative and qualitative analyses of intent narratives with respect to diversity, grounding, and plausibility; 3. benchmark tasks for AII and surrounding relevant contextual factors such as social involvement; 4. speech quality audio for all participants as well as privacy preserving multi-modal data, enabling lexical and nonverbal behavior analysis; and 5. coupling of self-reported goals of each participant (30 minute to 3 hour) with annotated AII (seconds).
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
V . Agrawal, A. Akinyemi, K. Alvero, M. Behrooz, J. Buffalini, F. M. Carlucci, J. Chen, J. Chen, Z. Chen, S. Cheng, et al. Seamless interaction: Dyadic audiovisual motion modeling and large-scale dataset.arXiv preprint arXiv:2506.22554, 2025
arXiv 2025
-
[2]
Akata, D
Z. Akata, D. Balliet, M. De Rijke, F. Dignum, V . Dignum, G. Eiben, A. Fokkens, D. Grossi, K. Hindriks, H. Hoos, et al. A research agenda for hybrid intelligence: augmenting human intel- lect with collaborative, adaptive, responsible, and explainable artificial intelligence.Computer, 53(8):18–28, 2020
2020
-
[3]
X. Alameda-Pineda, J. Staiano, R. Subramanian, L. M. Batrinca, E. Ricci, B. Lepri, O. Lanz, and N. Sebe. SALSA: A Novel Dataset for Multimodal Group Behavior Analysis.CoRR, abs/1506.06882, 2015. URLhttp://arxiv.org/abs/1506.06882
arXiv 2015
-
[4]
M. C. Ashton and K. Lee. The hexaco–60: A short measure of the major dimensions of personality.Journal of personality assessment, 91(4):340–345, 2009
2009
-
[5]
Z. Bai, P. Wang, T. Xiao, T. He, Z. Han, Z. Zhang, and M. Z. Shou. Hallucination of multimodal large language models: A survey.arXiv preprint arXiv:2404.18930, 2024
arXiv 2024
-
[6]
M. Bain, J. Huh, T. Han, and A. Zisserman. Whisperx: Time-accurate speech transcription of long-form audio.arXiv preprint arXiv:2303.00747, 2023
arXiv 2023
-
[7]
Belardinelli
A. Belardinelli. Gaze-based intention estimation: principles, methodologies, and applications in hri, 2023
2023
-
[8]
Bratman.Intention, plans, and practical reason
M. Bratman.Intention, plans, and practical reason. Harvard University Press, Cambridge, MA, 1987
1987
Show all 93 references
-
[9]
Broekens and W.-P
J. Broekens and W.-P. Brinkman. Affectbutton: A method for reliable and valid affective self-report.International Journal of Human-Computer Studies, 71(6):641–667, 2013
2013
-
[10]
Cabitza, A
F. Cabitza, A. Campagner, and V . Basile. Toward a perspectivist turn in ground truthing for predictive computing.Proceedings of the AAAI Conference on Artificial Intelligence, 37:6860–6868, 6 2023. ISSN 2374-3468. doi: 10.1609/AAAI.V37I6.25840. URL https: //ojs.aaai.org/index...
2023 doi
-
[11]
Cabrera-Quiros, A
L. Cabrera-Quiros, A. Demetriou, E. Gedik, L. van der Meij, and H. Hung. The matchnmingle dataset: a novel multi-sensor resource for the analysis of social interactions and group dynamics in-the-wild during free-standing conversations and speed dates.IEEE Transactions on Affec...
2018
-
[12]
Cabrera-Quiros, A
L. Cabrera-Quiros, A. Demetriou, E. Gedik, L. van der Meij, and H. Hung. The matchnmingle dataset: A novel multi-sensor resource for the analysis of social interactions and group dynamics in-the-wild during free-standing conversations and speed dates.IEEE Transactions on Affec...
2021
-
[13]
Calib.io calibrator
calib.io. Calib.io calibrator. URL https://calib.io/products/calib. Accessed: 2026- 04-22
2026
-
[14]
Carion, L
N. Carion, L. Gustafson, Y .-T. Hu, S. Debnath, R. Hu, D. Suris, C. Ryali, K. V . Alwala, H. Khedr, A. Huang, et al. Sam 3: Segment anything with concepts.arXiv preprint arXiv:2511.16719, 2025
2025 arXiv
-
[15]
J. M. Cheek and A. H. Buss. Shyness and sociability.Journal of personality and social psychology, 41(2):330, 1981
1981
-
[16]
K. Chen, Z. Zhang, W. Zeng, R. Zhang, F. Zhu, and R. Zhao. Shikra: Unleashing Multimodal LLM’s Referential Dialogue Magic, July 2023. URL http://arxiv.org/abs/2306.15195. arXiv:2306.15195 [cs]
2023 arXiv
-
[17]
K. M. Connor, J. R. Davidson, L. E. Churchill, A. Sherwood, R. H. Weisler, and E. Foa. Psychometric properties of the social phobia inventory (spin): New self-rating scale.The British Journal of Psychiatry, 176(4):379–386, 2000
2000
-
[18]
Dudzik and J
B. Dudzik and J. Broekens. A valid self-report is never late, nor is it early: On considering the "right" temporal distance for assessing emotional experience, 2023
2023
-
[19]
Dudzik, M.-P
B. Dudzik, M.-P. Jansen, F. Burger, F. Kaptein, J. Broekens, D. K. Heylen, H. Hung, M. A. Neerincx, and K. P. Truong. Context in human emotion perception for automatic affect detection: A survey of audiovisual databases. In2019 8th International Conference on Affective Computi...
2019
-
[20]
C. Edelsky. Who’s Got the Floor?Language in Society, 10(3):383–421, 1981. ISSN 0047-4045
1981
-
[21]
F. F, K. I, S. K, and W. C. Toward a script theory of guidance in computer-supported collaborative learning.Educ Psychol, 2013. doi: 10.1080/00461520.2012.748005
2013
-
[22]
Fonagy, P
P. Fonagy, P. Luyten, A. Moulton-Perkins, Y .-W. Lee, F. Warren, S. Howard, R. Ghinai, P. Fearon, and B. Lowyck. Development and validation of a self-report measure of mentalizing: The reflective functioning questionnaire.PloS one, 11(7):e0158678, 2016
2016
-
[23]
Gebru, J
T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. D. Iii, and K. Crawford. Datasheets for datasets.Communications of the ACM, 64(12):86–92, 2021
2021
-
[24]
Grauman, A
K. Grauman, A. Westbury, L. Torresani, K. Kitani, J. Malik, T. Afouras, K. Ashutosh, V . Baiyya, S. Bansal, B. Boote, et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. In2024 IEEE/CVF Conference on Computer Vision and Pattern Reco...
2024
-
[25]
J. Gu, X. Jiang, Z. Shi, H. Tan, X. Zhai, C. Xu, W. Li, Y . Shen, S. Ma, H. Liu, S. Wang, K. Zhang, Y . Wang, W. Gao, L. Ni, and J. Guo. A survey on llm-as-a-judge, 2025. URL https://arxiv.org/abs/2411.15594
2025 arXiv
-
[26]
Honnibal, I
M. Honnibal, I. Montani, S. Van Landeghem, A. Boyd, et al. spacy: Industrial-strength natural language processing in python. 2020
2020
-
[27]
T. M. Hrkalovic, B. Dudzik, D. Balliet, and H. Hung. Parsel: A multimodal dataset for modeling decision-making processes involved in selecting partners for joint tasks.IEEE Transactions on Affective Computing, 2025
2025
-
[28]
Hung and B
H. Hung and B. Kröse. Detecting F-formations as dominant sets. InProceedings of the 13th international conference on multimodal interfaces, pages 231–238, Alicante Spain, Nov. 2011. ACM. ISBN 978-1-4503-0641-6. doi: 10.1145/2070481.2070525. URL https://dl.acm. org/doi/10.1145/...
2011
-
[29]
H. Hung, L. Li, J. Molhoek, and J. Zhou. The discontent with intent estimation in-the-wild: the case for unrealized intentions. InExtended abstracts of the CHI conference on human factors in computing systems, pages 1–9, 2024
2024
-
[30]
life of the party
P. Ingram and M. W. Morris. Do people mix at mixers? structure, homophily, and the “life of the party”.Administrative Science Quarterly, 52(4):558–585, 2007
2007
-
[31]
T. Jing, T. Chen, R. Tian, Y . Chen, J. Domeyer, H. Toyoda, R. Sherony, and Z. Ding. Psi: A benchmark for human interpretation and response in traffic interactions. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2025
2025
-
[32]
Kendon.Conducting Interaction: Patterns of Behavior in Focused Encounters
A. Kendon.Conducting Interaction: Patterns of Behavior in Focused Encounters. Cambridge University Press, Cambridge, UK, 1990
1990
-
[33]
Khindkar, V
V . Khindkar, V . Balasubramanian, C. Arora, A. Subramanian, and C. Jawahar. Can reasons help improve pedestrian intent estimation? a cross-modal approach. In2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 11515–11522. IEEE, 2024
2024
-
[34]
A. W. Kruglanski, M. Chernikova, M. Babush, M. Dugas, and B. M. Schumpe. Chapter three - the architecture of goal systems: Multifinality, equifinality, and counterfinality in means—end relations. volume 2 ofAdvances in Motivation Science, pages 69–98. Elsevier, 2015. doi: 10.1...
2015 doi
-
[35]
Kuchaiev, J
O. Kuchaiev, J. Li, H. Nguyen, O. Hrinchuk, R. Leary, B. Ginsburg, S. Kriman, S. Beliaev, V . Lavrukhin, J. Cook, P. Castonguay, M. Popova, J. Huang, and J. M. Cohen. Nemo: a toolkit for building ai applications using neural modules, 2019
2019
-
[36]
R. D. Lennox and R. N. Wolfe. Revision of the self-monitoring scale. 1984
1984
-
[37]
J. Li, P. Wei, W. Han, and L. Fan. Intentqa: Context-aware video intent reasoning. InProceedings of the IEEE/CVF international conference on computer vision, pages 11963–11974, 2023
2023
-
[38]
Z. Lian, H. Chen, L. Chen, H. Sun, L. Sun, Y . Ren, Z. Cheng, B. Liu, R. Liu, X. Peng, J. Yi, and J. Tao. AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models, May 2025. URL http://arxiv.org/abs/2501. 16566. arXiv:2501....
2025 arXiv
-
[39]
Martín-Fernández, B
M. Martín-Fernández, B. Requero, X. Zhou, D. Gonçalves, and D. Santos. Refinement of the analysis-holism scale: A cross-cultural adaptation and validation of two shortened measures of analytic versus holistic thinking in spain and the united states.Personality and Individual D...
2022
-
[40]
Mathur, M
L. Mathur, M. Qian, P. P. Liang, and L.-P. Morency. Social Genome: Grounded Social Reasoning Abilities of Multimodal Models, Feb. 2025. URL http://arxiv.org/abs/2502.15109. arXiv:2502.15109 [cs]
2025 arXiv
-
[41]
A. R. Mele.Springs of action: Understanding intentional behavior. Oxford University Press, New York, USA, 1992
1992
-
[42]
Presidio: Data Protection and De-identification SDK
Microsoft. Presidio: Data Protection and De-identification SDK. https://microsoft. github.io/presidio/
-
[43]
V . S. R. Parupudi. Magnitude matters: a superior class of similarity metrics for holistic semantic understanding, 2025. URLhttps://arxiv.org/abs/2509.19323
2025
-
[44]
Z. Peng, W. Wang, L. Dong, Y . Hao, S. Huang, S. Ma, and F. Wei. Kosmos-2: Grounding Multimodal Large Language Models to the World, July 2023. URL http://arxiv.org/abs/ 2306.14824. arXiv:2306.14824 [cs]
2023 arXiv
-
[45]
J. V . Quiros, C. Raman, S. Tan, E. Gedik, L. Cabrera-Quiros, and H. Hung. Rewind dataset: Privacy-preserving speaking status segmentation from multimodal body movement signals in the wild, 2024. 12
2024
-
[46]
Raman, J
C. Raman, J. Vargas Quiros, S. Tan, A. Islam, E. Gedik, and H. Hung. Conflab: A data collection concept, dataset, and benchmark for machine analysis of free-standing social interactions in the wild. In S. Koyejo, S. Mohamed, A. Agarwal, D. Bel- grave, K. Cho, and A. Oh, editor...
2022
-
[47]
Rasouli, I
A. Rasouli, I. Kotseruba, T. Kunic, and J. K. Tsotsos. Pie: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction. InInternational Conference on Computer Vision (ICCV), pages 6262–6271, Seoul, South Korea, 2019. IEEE
2019
-
[48]
J. F. Rauthmann and R. A. Sherman. Chapter 13 - conceptualizing and measuring the psychological situation. In D. Wood, S. J. Read, P. Harms, and A. Slaughter, editors, Measuring and Modeling Persons and Situations, pages 427–463. Academic Press, 2021. ISBN 978-0-12-819200-9. d...
2021 doi
-
[49]
Rme 32 ad,
RME. Rme 32 ad, . URL https://rme-audio.de/m-32-m-16-ad.html . Accessed: 2026- 04-22
2026
-
[50]
Rme fireface ufx iii,
RME. Rme fireface ufx iii, . URL https://rme-audio.de/fireface-ufx-3.html . Ac- cessed: 2026-04-22
2026
-
[51]
R. B. Rubin and M. M. Martin. Development of a measure of interpersonal communication competence.Communication Research Reports, 11(1):33–44, 1994
1994
-
[52]
Rudenko, L
A. Rudenko, L. Palmieri, M. Herman, K. M. Kitani, D. M. Gavrila, and K. O. Arras. Human motion trajectory prediction: a survey.The International Journal of Robotics Research, 39 (8):895–935, 2020. doi: 10.1177/0278364920917446. URL https://doi.org/10.1177/ 0278364920917446
2020 doi
-
[53]
R. C. Schank and R. P. Abelson.Scripts, Plans, Goals, and Understanding: An Inquiry Into Human Knowledge Structures. Psychology Press, New York, 1977. ISBN 9780203781036. doi: 10.4324/9780203781036
1977 doi
-
[54]
Sennheiser sk 20 bodypack transmitter xs wireless series,
Sennheiser. Sennheiser sk 20 bodypack transmitter xs wireless series, . URL https://docs. cloud.sennheiser.com/en-us/xsw/xsw/manual-sk-overview.html . Accessed: 2026- 04-22
2026
-
[55]
Sennheiser sk 2000 bodypack transmitter,
Sennheiser. Sennheiser sk 2000 bodypack transmitter, . URL https://www.sennheiser. com/en-us/catalog/products/wireless-systems/sk-2000. Accessed: 2026-04-22
2000
-
[56]
Setti, C
F. Setti, C. Russell, C. Bassetti, and M. Cristani. F-formation detection: Individuating free- standing conversational groups in images.PLoS ONE, 10, 2015
2015
-
[57]
Vista omni anchor,
Sewio. Vista omni anchor, . URL https://docs.sewio.net/docs/ anchor-vista-omni-30147663.html. Accessed: 2026-04-22
2026
-
[58]
Leonardo personal tag,
Sewio. Leonardo personal tag, . URL https://docs.sewio.net/docs/ tag-leonardo-personal-30146967.html. Accessed: 2026-04-22
2026
-
[59]
Swofford, J
M. Swofford, J. Peruzzi, N. Tsoi, S. Thompson, R. Martín-Martín, S. Savarese, and M. Vázquez. Improving social awareness through dante: Deep affinity network for clustering conversational interactants.Proceedings of the ACM on Human-Computer Interaction, 4(CSCW1):1–23, 2020
2020
-
[60]
S. Tan, D. M. Tax, and H. Hung. Conversation group detection with spatio-temporal con- text. InProceedings of the 2022 International Conference on Multimodal Interaction, ICMI ’22, page 170–180, New York, NY , USA, 2022. Association for Computing Machinery. ISBN 9781450393904....
2022
-
[61]
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295, 2024
2024 arXiv
-
[62]
Umagami, L
R. Umagami, L. Yue, X. Chu, R. Fukushima, T. Narita, Y . Mukuta, T. Takahata, J. Yang, and T. Harada. Intend to move: A multimodal dataset for intention-aware human motion understanding. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and...
-
[63]
Umagami, L
R. Umagami, L. Yue, X. Chu, R. Fukushima, T. Narita, Y . Mukuta, T. Takahata, J. Yang, and T. Harada. Intend to move: A multimodal dataset for intention-aware human motion understanding. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and...
2026
-
[64]
J. Vargas. Covfee: Continuous Video Feedback Tool. URL https://github.com/josedvq/ covfee. Accessed: 2021-05-28
2021
-
[65]
Vargas Quiros, S
J. Vargas Quiros, S. Tan, C. Raman, E. Gedik, I. Pronotaris, and H. Hung. Spcl mingle badge. URL https://github.com/TUDelft-SPC-Lab/spcl_midge_hardware . Accessed: 2026- 04-22
2026
-
[66]
Vascon, E
S. Vascon, E. Z. Mequanint, M. Cristani, H. Hung, M. Pelillo, and V . Murino. Detecting conversa- tional groups in images and sequences: A robust game-theoretic approach.Computer Vision and Image Understanding, 143:11–24, Feb. 2016. ISSN 1077-3142. doi: 10.1016/j.cviu.2015.09....
2016 doi
-
[67]
B. Vissa. Agency in action: Entrepreneurs’ networking style and initiation of economic exchange.Organization Science, 23(2):492–510, 2012
2012
-
[68]
Y . Wu, J. Xiong, and X. Deng. How Social is It? A Benchmark for LLMs’ Capabilities in Multi- user Multi-turn Social Agent Tasks, Apr. 2025. URL http://arxiv.org/abs/2505.04628. arXiv:2505.04628 [cs]
2025 arXiv
-
[69]
Y . Xu, J. Zhang, Q. Zhang, and D. Tao. Vitpose: Simple vision transformer baselines for human pose estimation.Advances in neural information processing systems, 35:38571–38584, 2022
2022
-
[70]
Zhang, H
H. Zhang, H. Xu, X. Wang, Q. Zhou, S. Zhao, and J. Teng. Mintrec: A new dataset for multimodal intent recognition. InProceedings of the 30th ACM International Conference on Multimedia, MM ’22, page 1688–1697. ACM, Oct. 2022. doi: 10.1145/3503161.3547906. URL http://dx.doi.org/...
2022
-
[71]
Zhang, X
H. Zhang, X. Wang, H. Xu, Q. Zhou, K. Gao, J. Su, jinyue Zhao, W. Li, and Y . Chen. MIntrec2.0: A large-scale benchmark dataset for multimodal intent recognition and out-of-scope detection in conversations. InThe Twelfth International Conference on Learning Representations, 20...
2024
-
[72]
Zhang, Z
H. Zhang, Z. Li, Y . Zhu, H. Xu, P. Wang, H. Zhu, J. Zhou, and J. Zhang. Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark, Apr
-
[73]
Zhang, P
S. Zhang, P. Sun, S. Chen, M. Xiao, W. Shao, W. Zhang, Y . Liu, K. Chen, and P. Luo. GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest. In A. Del Bue, C. Canton, J. Pont-Tuset, and T. Tommasi, editors,Computer Vision – ECCV 2024 Workshops, pages 52–70, Cha...
2024 doi
-
[74]
X. Zhou, H. Zhu, L. Mathur, R. Zhang, H. Yu, Z. Qi, L.-P. Morency, Y . Bisk, D. Fried, G. Neubig, and M. Sap. SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents, Mar
-
[78]
Do you see multiple possibilities? Click ‘‘+’’to add more entries if you see multiple intentions or interpretations
-
[79]
Use your first impression and intuition; there is no right or wrong answer
Give your reasoning Complete the questionnaire under the video. Use your first impression and intuition; there is no right or wrong answer
-
[80]
If you have annotated the full video, you may submit
Continue the search Resume the video and repeat this process for every new intention you see until the clip ends. If you have annotated the full video, you may submit. **Grading rubric:** Your response will be evaluated based on the following criteria: Intention (not just acti...
-
[81]
Adjust the timestamp to mark the exact start and end of the intention
Watch and pause Watch the clip and pause as soon as you notice an intention. Adjust the timestamp to mark the exact start and end of the intention. 24
-
[83]
No intention seen
Give your reasoning Complete the questionnaire under the video. Use your first impression and intuition; there is no right or wrong answer. Intention (Free text): Confidence: On a scale of 1-5, how confident are you in this interpretation? Just a guess / Extremely confident [L...
-
[84]
Adjust the timestamp to mark the exact start and end of the intention
Watch and pause Watch the clip and pause as soon as you notice an intention. Adjust the timestamp to mark the exact start and end of the intention
-
[85]
Do you see multiple possibilities? Click ‘‘+’’ to add more entries if you see multiple intentions or interpretations
-
[86]
No intention seen
Give your reasoning Complete the questionnaire under the video. Use your first impression and intuition; there is no right or wrong answer. Intention (Free text): Confidence: On a scale of 1-5, how confident are you in this interpretation? Just a guess / Extremely confident [L...
-
[87]
Choose one ‘initiative_type‘
-
[88]
yeah", "okay
Choose one ‘uptake_status‘. A topic-redirection attempt is a turn that tries to shift, reopen, broaden, or redirect the group’s shared attention to a different local topical line. Important: Judge whether the marked turn attempts a redirection based on what it does at that mom...
-
[89]
Timestamps: [start time] [end time]
-
[90]
Intention: [brief description of what the participant is trying to do or achieve]
-
[91]
Confidence in interpretation: [1-5]
-
[92]
Why: [observable cues plus your interpretation of the participant’s beliefs/desires and the situation]
-
[93]
Confidence in explanation: [1-5]
-
[94]
Intensity: [1-5 rating of how strong or high-priority the intention appears]
-
[95]
who is talking with whom
Counterfactual explanation: [an alternative interpretation that would change the understood intention] If the full clip contains no clear intention, answer only: No clear intention: yes Why: [brief explanation] In the above prompts, <image> refers to the image of that particul...
-
[2024]
no intention found
URLhttp://arxiv.org/abs/2310.11667. arXiv:2310.11667 [cs]. 14 COSI-Lab Appendices The Appendices sections are inspired by the Datasheet for Datasets [23] and are organized as follows: A Dataset documentation Motivation and Intended UseCOSI-Lab is a multimodal multisensor datas...
1920 arXiv
- [2025]
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.