REVIEW 4 major objections 6 minor 93 references
This paper argues that perceived social intentions are inherently multiple and perspective-dependent, and presents COSI-Lab, a dataset that aligns participants' long-term goals with second-level, multi-observer 'apparent intention' annotati
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 00:48 UTC pith:C22Q3NW2
load-bearing objection A genuinely useful dataset resource for social-intention research, but its core AII annotations are unvalidated, so the headline coupling claim is a promise rather than a demonstrated result. the 4 major comments →
COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper argues that in weakly scripted social settings, a person's apparent intention is not a single ground-truth label but a space of plausible interpretations that different perceivers construct from observable cues, situation assumptions, and their own interpretative tendencies. COSI-Lab operationalizes this by having crowd-sourced annotators watch 30-second overhead video clips of real mingling and write free-form intention narratives with timestamps, confidence, intensity, and counterfactual alternatives, while also capturing annotator traits and demographics. The resulting dataset is claimed to be the first to align these multi-perspective apparent-intention annotations with the obs
What carries the argument
The Apparent Intent Inference (AII) problem: the task of inferring, from an ex-situ third-person perspective, what intention a person appears to have at a given moment, independent of whether that intention is later realized. The annotation protocol is the load-bearing instrument: it structures narratives through cues, situation characteristics, situation classes, and social scripts, and adds annotator-level trait data so that multiplicity of interpretations can be studied as a function of the perceiver rather than as noise.
Load-bearing premise
The claim depends on ex-situ online annotators' free-form narratives being a valid window into genuine variation in intention perception, rather than arbitrary responses produced by poorly motivated crowd workers.
What would settle it
Show that annotators' perceived intentions do not track any behavioral or self-report signal, for instance that participants' post-session goal-attainment ratings share no systematic association with the apparent intentions narrated by observers, or that two annotators' narratives for the same clip are no more similar to each other than to a shuffled baseline. More directly, if participants had reported their immediate proximal intentions during the event and observer narratives matched those reports at chance, the resource's validity as intention-perception data would collapse.
If this is right
- Researchers can study the relationship between self-reported long-term goals and apparent proximal intentions on the same individuals at the same event for the first time.
- Intention-aware systems can be trained or evaluated to output multiple plausible intention hypotheses with explanations, rather than being forced to commit to a single label.
- The dataset provides benchmark tasks for apparent intent inference and conversation group detection, including a human-LLM comparison that reveals systematic differences in how people and models describe intentions.
- Multi-perspective annotations coupled with annotator traits allow investigation of how demographics and reflective functioning shape intention perception, supporting perspectivist modeling.
- The weakly scripted, ecologically valid setting with real professional consequences makes findings more likely to transfer to in-the-wild social interactions than strongly scripted datasets.
Where Pith is reading between the lines
- If the perspective-driven framing is correct, agreement between annotators should not be the primary quality metric; instead, downstream systems could be evaluated on the diversity and plausibility of the interpretations they generate, a departure from typical label-agreement benchmarks.
- The dataset could be extended with participants' own in-the-moment proximal intention self-reports, for example via brief experience sampling between sessions, to directly validate whether ex-situ observer narratives track the observed person's actual immediate intentions.
- The goal-to-perceived-intention link opens a route to study unrealized intentions: cases where observers perceive a goal-directed intention that the participant later reports not achieving, a key gap the paper identifies in existing intention research.
- Because annotator traits and demographics are collected, the data could support deliberately resampling perspectives to make intention-inference systems aware of who is perceiving, potentially reducing demographic bias in social AI.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. COSI-Lab is a multimodal dataset of a two-session conference mingling event involving 32 academics, with overhead video, close-talk and wearable audio, IMU, UWB, surveys, and annotations. The paper's main contribution is an annotation protocol that elicits multi-perspective 'apparent intention' narratives from crowd annotators alongside self-reported participant goals, and the claim that this is the first resource coupling long-term goals with short-term third-party perceived intentions in a weakly scripted social setting. The paper also presents descriptive analyses of conversation groups, topics, and goals, plus benchmark demonstrations for AII and conversation-group detection.
Significance. If the AII annotations are meaningful, the dataset is a potentially valuable resource for perspectivist, explainable intention-inference research: it includes synchronized multimodal data, privacy-preserving processing, an annotation protocol with theory-grounded components (cues, assumptions, scripts), and an honest LLM-as-judge evaluation that reports a 45.3% discrimination accuracy. The detailed data collection and processing description, camera-wise 5-fold group detection evaluation, and released code for reproducibility are strengths. However, the core value rests on the validity of the AII narratives, and that validity is not yet demonstrated.
major comments (4)
- [§5.1, §7.1, §8 (Cue Grounding)] The load-bearing claim—that COSI-Lab couples self-reported long-term goals with multi-perspective apparent-intention annotations—is not supported by any reliability or validity evidence for the AII annotations. There is no inter-annotator agreement metric, no verification that cues cited in the narratives are present in the video/audio, and no analysis relating the AII narratives to the observed participants' own self-reported goals. Section 8 states that cue grounding 'was not assessed' and that goals and AII 'bare some relationship' without testing it. This makes the central dataset contribution, and desiderata D1–D4 built on it, currently an unverified assertion. At minimum, the paper should report inter-annotator agreement (e.g., Krippendorff's alpha on intention categories or annotation components), a cue-presence check on a sample, and an exploratory comparison between AII narrativ
- [§6, Appendix E.2/E.3] The quantitative descriptive claims in Figure 4 (topic composition and topic-redirection uptake) are derived entirely from LLM-assisted coding with no human validation. The appendix provides detailed prompts, but no inter-coder reliability against human coders, and no two-human agreement baseline. As reported, these statistics are trustworthy only insofar as the Qwen3-14B labels are reliable, which is not established. The authors should validate a sample of the LLM coding against human coders and report agreement; otherwise these results should be presented as illustrative rather than quantitative findings.
- [§7.1] The 'intention estimation' benchmark is currently descriptive, not predictive. It compares human and Gemma4-generated narratives via semantic similarity and an LLM-as-judge accuracy of 45.3%, but no concrete task formulation, evaluation metric, or baseline protocol is given that future models could be measured against. The 45.3% figure is reported without variance or confidence intervals and without a significance test, so its interpretation is unclear. To claim a benchmark task, the paper should define an evaluation setup (e.g., predicting annotation components, matching a target perspective, or ranking narrative plausibility) and provide a reproducible metric.
- [§1, Table 1, §9] The claim of satisfying desideratum D4 ('hierarchical structure of human intention') and the conclusion that COSI-Lab 'for the first time' enables study of 'subgoals or apparent proximal intentions' go beyond what is demonstrated. The dataset contains self-reported long-term goals and AII for short clips, but no annotation explicitly linking proximal intentions to subgoals or to distal goal hierarchies. The checkmark for D4 in Table 1 and the conclusion wording should be softened or accompanied by an explicit description of which hierarchy-related variables are present in the released data.
minor comments (6)
- [Abstract] 'a explainable' should be 'an explainable'.
- [§5.1] 'the begging and end of the intention' should be 'the beginning and end'.
- [§8] 'they bare some relationship' should be 'they bear some relationship'.
- [Appendix C.2.1] 'felxible inputs' should be 'flexible inputs'.
- [Appendix B] 'the michrophones record' should be 'the microphones record'.
- [§7.1 / Figure 5] The semantic-similarity plots are described qualitatively as showing 'clear clustering separation'; a quantitative cluster-separation or distance statistic would make the claim more precise.
Circularity Check
No significant circularity: COSI-Lab is a resource paper whose benchmarks and analyses are self-contained or explicitly deferred.
full rationale
The paper's central contribution is the COSI-Lab dataset itself—a multimodal recording of a weakly scripted workshop with self-reported goals, multi-perspective AII annotations, and benchmark tasks. There is no derivation chain in which a predicted quantity is constructed from the same quantity it claims to predict, and no fitted parameter is renamed as a prediction. The AII annotations are elicited from crowd annotators and then compared with LLM outputs; the LLM-as-judge result (45.3%) is reported as a failure to distinguish human from model narratives, not as evidence for model superiority. The conversation-group detection benchmark trains standard models (DANTE, LSTM) on position/orientation features and evaluates against human group annotations; this is a conventional supervised evaluation, not circular. The descriptive topic-redirection analyses do rely on LLM-generated labels without human validation, which is a validity concern, but it is not circularity: the LLM is not being used to 'predict' a label that was itself derived from the LLM in a way that defines the result. The paper also openly concedes limitations that would be relevant to correctness or validity but not to circularity: cue grounding is not assessed, and the relationship between self-reported goals and AII is asserted rather than tested. Self-citations to ConfLab [46] and related prior work place the dataset in a lineage but do not function as an unverified load-bearing premise for the paper's conclusions. Overall, no step reduces to its own inputs by construction, so the circularity score is 0.
Axiom & Free-Parameter Ledger
free parameters (4)
- Assumed body height H =
1.7 m
- Keypoint height ratios α_i =
from COCO model (not given)
- LLM coding window parameters =
300-turn window, 250-turn stride (topic distribution); 300/220/80 for redirection screening
- Frame strides for baselines =
20 (LSTM), 300 (DANTE)
axioms (6)
- domain assumption The workshop setting is ecologically valid for studying social intentions despite being recorded and weakly scripted.
- domain assumption Head orientation is a valid indicator of conversational attention and group membership (Kendon's F-formation).
- domain assumption 30-second clips are sufficient for a third-party observer to perceive proximal intentions.
- ad hoc to paper LLM-generated topic/redirection labels are acceptably reliable for descriptive statistics without human-validation metrics.
- domain assumption If multiple plausible intention narratives are socially meaningful, evaluating them requires diversity/grounding/plausibility measures.
- standard math Hyperbolic Tangent Similarity (HTS) is a valid semantic similarity metric.
Cite this review
Pith. "Pith review of COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention." pith.science (2026). https://pith.science/paper/C22Q3NW2
@misc{pith2026260728649,
author = {Pith},
title = {Pith review of: COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention},
year = {2026},
howpublished = {\url{https://pith.science/paper/C22Q3NW2}},
note = {Machine review of arXiv:2607.28649}
}
read the original abstract
COSI-Lab presents a multimodal, multi-sensor dataset of an interdisciplinary scientific workshop containing 32 academics at an international conference. It captures ecologically valid social interactions in a weakly scripted setting consisting of two 30-minute mingling sessions with real professional and social consequences for the participants involved. We argue that future intelligent systems could be better equipped to handle subjective perceptions by modeling their multiplicity not as label noise but as a explainable perspective-driven reasoning process. We focus on the Apparent Intent Inference (AII) problem as determined by ex-situ observers and conceptualize intentions to be independent of manifest future outcomes. We contribute 1. a novel annotation process for AII that accounts for a perceiver's own interpretative tendencies, 2. quantitative and qualitative analyses of intent narratives with respect to diversity, grounding, and plausibility; 3. benchmark tasks for AII and surrounding relevant contextual factors such as social involvement; 4. speech quality audio for all participants as well as privacy preserving multi-modal data, enabling lexical and nonverbal behavior analysis; and 5. coupling of self-reported goals of each participant (30 minute to 3 hour) with annotated AII (seconds).
Figures
Reference graph
Works this paper leans on
-
[1]
V . Agrawal, A. Akinyemi, K. Alvero, M. Behrooz, J. Buffalini, F. M. Carlucci, J. Chen, J. Chen, Z. Chen, S. Cheng, et al. Seamless interaction: Dyadic audiovisual motion modeling and large-scale dataset.arXiv preprint arXiv:2506.22554, 2025
Pith/arXiv arXiv 2025
-
[2]
Akata, D
Z. Akata, D. Balliet, M. De Rijke, F. Dignum, V . Dignum, G. Eiben, A. Fokkens, D. Grossi, K. Hindriks, H. Hoos, et al. A research agenda for hybrid intelligence: augmenting human intel- lect with collaborative, adaptive, responsible, and explainable artificial intelligence.Computer, 53(8):18–28, 2020
2020
-
[3]
X. Alameda-Pineda, J. Staiano, R. Subramanian, L. M. Batrinca, E. Ricci, B. Lepri, O. Lanz, and N. Sebe. SALSA: A Novel Dataset for Multimodal Group Behavior Analysis.CoRR, abs/1506.06882, 2015. URLhttp://arxiv.org/abs/1506.06882
Pith/arXiv arXiv 2015
-
[4]
M. C. Ashton and K. Lee. The hexaco–60: A short measure of the major dimensions of personality.Journal of personality assessment, 91(4):340–345, 2009
2009
-
[5]
Z. Bai, P. Wang, T. Xiao, T. He, Z. Han, Z. Zhang, and M. Z. Shou. Hallucination of multimodal large language models: A survey.arXiv preprint arXiv:2404.18930, 2024
Pith/arXiv arXiv 2024
-
[6]
M. Bain, J. Huh, T. Han, and A. Zisserman. Whisperx: Time-accurate speech transcription of long-form audio.arXiv preprint arXiv:2303.00747, 2023
Pith/arXiv arXiv 2023
-
[7]
Belardinelli
A. Belardinelli. Gaze-based intention estimation: principles, methodologies, and applications in hri, 2023
2023
-
[8]
Bratman.Intention, plans, and practical reason
M. Bratman.Intention, plans, and practical reason. Harvard University Press, Cambridge, MA, 1987
1987
-
[9]
Broekens and W.-P
J. Broekens and W.-P. Brinkman. Affectbutton: A method for reliable and valid affective self-report.International Journal of Human-Computer Studies, 71(6):641–667, 2013
2013
-
[10]
F. Cabitza, A. Campagner, and V . Basile. Toward a perspectivist turn in ground truthing for predictive computing.Proceedings of the AAAI Conference on Artificial Intelligence, 37:6860–6868, 6 2023. ISSN 2374-3468. doi: 10.1609/AAAI.V37I6.25840. URL https: //ojs.aaai.org/index.php/AAAI/article/view/25840
-
[11]
Cabrera-Quiros, A
L. Cabrera-Quiros, A. Demetriou, E. Gedik, L. van der Meij, and H. Hung. The matchnmingle dataset: a novel multi-sensor resource for the analysis of social interactions and group dynamics in-the-wild during free-standing conversations and speed dates.IEEE Transactions on Affective Computing, 12(1):113–130, 2018. 10
2018
-
[12]
Cabrera-Quiros, A
L. Cabrera-Quiros, A. Demetriou, E. Gedik, L. van der Meij, and H. Hung. The matchnmingle dataset: A novel multi-sensor resource for the analysis of social interactions and group dynamics in-the-wild during free-standing conversations and speed dates.IEEE Transactions on Affective Computing, 12(1):113–130, 2021
2021
-
[13]
Calib.io calibrator
calib.io. Calib.io calibrator. URL https://calib.io/products/calib. Accessed: 2026- 04-22
2026
-
[14]
N. Carion, L. Gustafson, Y .-T. Hu, S. Debnath, R. Hu, D. Suris, C. Ryali, K. V . Alwala, H. Khedr, A. Huang, et al. Sam 3: Segment anything with concepts.arXiv preprint arXiv:2511.16719, 2025
Pith/arXiv arXiv 2025
-
[15]
J. M. Cheek and A. H. Buss. Shyness and sociability.Journal of personality and social psychology, 41(2):330, 1981
1981
-
[16]
K. Chen, Z. Zhang, W. Zeng, R. Zhang, F. Zhu, and R. Zhao. Shikra: Unleashing Multimodal LLM’s Referential Dialogue Magic, July 2023. URL http://arxiv.org/abs/2306.15195. arXiv:2306.15195 [cs]
Pith/arXiv arXiv 2023
-
[17]
K. M. Connor, J. R. Davidson, L. E. Churchill, A. Sherwood, R. H. Weisler, and E. Foa. Psychometric properties of the social phobia inventory (spin): New self-rating scale.The British Journal of Psychiatry, 176(4):379–386, 2000
2000
-
[18]
Dudzik and J
B. Dudzik and J. Broekens. A valid self-report is never late, nor is it early: On considering the "right" temporal distance for assessing emotional experience, 2023
2023
-
[19]
Dudzik, M.-P
B. Dudzik, M.-P. Jansen, F. Burger, F. Kaptein, J. Broekens, D. K. Heylen, H. Hung, M. A. Neerincx, and K. P. Truong. Context in human emotion perception for automatic affect detection: A survey of audiovisual databases. In2019 8th International Conference on Affective Computing and Intelligent Interaction (ACII), pages 206–212, Cambridge, UK, 2019. IEEE
2019
-
[20]
C. Edelsky. Who’s Got the Floor?Language in Society, 10(3):383–421, 1981. ISSN 0047-4045
1981
-
[21]
F. F, K. I, S. K, and W. C. Toward a script theory of guidance in computer-supported collaborative learning.Educ Psychol, 2013. doi: 10.1080/00461520.2012.748005
arXiv 2013
-
[22]
Fonagy, P
P. Fonagy, P. Luyten, A. Moulton-Perkins, Y .-W. Lee, F. Warren, S. Howard, R. Ghinai, P. Fearon, and B. Lowyck. Development and validation of a self-report measure of mentalizing: The reflective functioning questionnaire.PloS one, 11(7):e0158678, 2016
2016
-
[23]
Gebru, J
T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. D. Iii, and K. Crawford. Datasheets for datasets.Communications of the ACM, 64(12):86–92, 2021
2021
-
[24]
Grauman, A
K. Grauman, A. Westbury, L. Torresani, K. Kitani, J. Malik, T. Afouras, K. Ashutosh, V . Baiyya, S. Bansal, B. Boote, et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19383–19400. IEEE Computer Society, 2024
2024
-
[25]
J. Gu, X. Jiang, Z. Shi, H. Tan, X. Zhai, C. Xu, W. Li, Y . Shen, S. Ma, H. Liu, S. Wang, K. Zhang, Y . Wang, W. Gao, L. Ni, and J. Guo. A survey on llm-as-a-judge, 2025. URL https://arxiv.org/abs/2411.15594
Pith/arXiv arXiv 2025
-
[26]
Honnibal, I
M. Honnibal, I. Montani, S. Van Landeghem, A. Boyd, et al. spacy: Industrial-strength natural language processing in python. 2020
2020
-
[27]
T. M. Hrkalovic, B. Dudzik, D. Balliet, and H. Hung. Parsel: A multimodal dataset for modeling decision-making processes involved in selecting partners for joint tasks.IEEE Transactions on Affective Computing, 2025
2025
-
[28]
H. Hung and B. Kröse. Detecting F-formations as dominant sets. InProceedings of the 13th international conference on multimodal interfaces, pages 231–238, Alicante Spain, Nov. 2011. ACM. ISBN 978-1-4503-0641-6. doi: 10.1145/2070481.2070525. URL https://dl.acm. org/doi/10.1145/2070481.2070525. 11
arXiv 2011
-
[29]
H. Hung, L. Li, J. Molhoek, and J. Zhou. The discontent with intent estimation in-the-wild: the case for unrealized intentions. InExtended abstracts of the CHI conference on human factors in computing systems, pages 1–9, 2024
2024
-
[30]
life of the party
P. Ingram and M. W. Morris. Do people mix at mixers? structure, homophily, and the “life of the party”.Administrative Science Quarterly, 52(4):558–585, 2007
2007
-
[31]
T. Jing, T. Chen, R. Tian, Y . Chen, J. Domeyer, H. Toyoda, R. Sherony, and Z. Ding. Psi: A benchmark for human interpretation and response in traffic interactions. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2025
2025
-
[32]
Kendon.Conducting Interaction: Patterns of Behavior in Focused Encounters
A. Kendon.Conducting Interaction: Patterns of Behavior in Focused Encounters. Cambridge University Press, Cambridge, UK, 1990
1990
-
[33]
Khindkar, V
V . Khindkar, V . Balasubramanian, C. Arora, A. Subramanian, and C. Jawahar. Can reasons help improve pedestrian intent estimation? a cross-modal approach. In2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 11515–11522. IEEE, 2024
2024
-
[34]
A. W. Kruglanski, M. Chernikova, M. Babush, M. Dugas, and B. M. Schumpe. Chapter three - the architecture of goal systems: Multifinality, equifinality, and counterfinality in means—end relations. volume 2 ofAdvances in Motivation Science, pages 69–98. Elsevier, 2015. doi: 10.1016/bs.adms.2015.04.001
-
[35]
Kuchaiev, J
O. Kuchaiev, J. Li, H. Nguyen, O. Hrinchuk, R. Leary, B. Ginsburg, S. Kriman, S. Beliaev, V . Lavrukhin, J. Cook, P. Castonguay, M. Popova, J. Huang, and J. M. Cohen. Nemo: a toolkit for building ai applications using neural modules, 2019
2019
-
[36]
R. D. Lennox and R. N. Wolfe. Revision of the self-monitoring scale. 1984
1984
-
[37]
J. Li, P. Wei, W. Han, and L. Fan. Intentqa: Context-aware video intent reasoning. InProceedings of the IEEE/CVF international conference on computer vision, pages 11963–11974, 2023
2023
-
[38]
Z. Lian, H. Chen, L. Chen, H. Sun, L. Sun, Y . Ren, Z. Cheng, B. Liu, R. Liu, X. Peng, J. Yi, and J. Tao. AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models, May 2025. URL http://arxiv.org/abs/2501. 16566. arXiv:2501.16566 [cs]
Pith/arXiv arXiv 2025
-
[39]
M. Martín-Fernández, B. Requero, X. Zhou, D. Gonçalves, and D. Santos. Refinement of the analysis-holism scale: A cross-cultural adaptation and validation of two shortened measures of analytic versus holistic thinking in spain and the united states.Personality and Individual Dif- ferences, 186:111322, 2022. ISSN 0191-8869. doi: https://doi.org/10.1016/j.p...
arXiv 2022
-
[40]
L. Mathur, M. Qian, P. P. Liang, and L.-P. Morency. Social Genome: Grounded Social Reasoning Abilities of Multimodal Models, Feb. 2025. URL http://arxiv.org/abs/2502.15109. arXiv:2502.15109 [cs]
Pith/arXiv arXiv 2025
-
[41]
A. R. Mele.Springs of action: Understanding intentional behavior. Oxford University Press, New York, USA, 1992
1992
-
[42]
Presidio: Data Protection and De-identification SDK
Microsoft. Presidio: Data Protection and De-identification SDK. https://microsoft. github.io/presidio/
-
[43]
V . S. R. Parupudi. Magnitude matters: a superior class of similarity metrics for holistic semantic understanding, 2025. URLhttps://arxiv.org/abs/2509.19323
arXiv 2025
-
[44]
Z. Peng, W. Wang, L. Dong, Y . Hao, S. Huang, S. Ma, and F. Wei. Kosmos-2: Grounding Multimodal Large Language Models to the World, July 2023. URL http://arxiv.org/abs/ 2306.14824. arXiv:2306.14824 [cs]
Pith/arXiv arXiv 2023
-
[45]
J. V . Quiros, C. Raman, S. Tan, E. Gedik, L. Cabrera-Quiros, and H. Hung. Rewind dataset: Privacy-preserving speaking status segmentation from multimodal body movement signals in the wild, 2024. 12
2024
-
[46]
Raman, J
C. Raman, J. Vargas Quiros, S. Tan, A. Islam, E. Gedik, and H. Hung. Conflab: A data collection concept, dataset, and benchmark for machine analysis of free-standing social interactions in the wild. In S. Koyejo, S. Mohamed, A. Agarwal, D. Bel- grave, K. Cho, and A. Oh, editors,Advances in Neural Information Processing Sys- tems, volume 35, pages 23701–23...
2022
-
[47]
Rasouli, I
A. Rasouli, I. Kotseruba, T. Kunic, and J. K. Tsotsos. Pie: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction. InInternational Conference on Computer Vision (ICCV), pages 6262–6271, Seoul, South Korea, 2019. IEEE
2019
-
[48]
J. F. Rauthmann and R. A. Sherman. Chapter 13 - conceptualizing and measuring the psychological situation. In D. Wood, S. J. Read, P. Harms, and A. Slaughter, editors, Measuring and Modeling Persons and Situations, pages 427–463. Academic Press, 2021. ISBN 978-0-12-819200-9. doi: https://doi.org/10.1016/B978-0-12-819200-9.00009-0. URL https://www.scienced...
-
[49]
Rme 32 ad,
RME. Rme 32 ad, . URL https://rme-audio.de/m-32-m-16-ad.html . Accessed: 2026- 04-22
2026
-
[50]
Rme fireface ufx iii,
RME. Rme fireface ufx iii, . URL https://rme-audio.de/fireface-ufx-3.html . Ac- cessed: 2026-04-22
2026
-
[51]
R. B. Rubin and M. M. Martin. Development of a measure of interpersonal communication competence.Communication Research Reports, 11(1):33–44, 1994
1994
-
[52]
A. Rudenko, L. Palmieri, M. Herman, K. M. Kitani, D. M. Gavrila, and K. O. Arras. Human motion trajectory prediction: a survey.The International Journal of Robotics Research, 39 (8):895–935, 2020. doi: 10.1177/0278364920917446. URL https://doi.org/10.1177/ 0278364920917446
-
[53]
R. C. Schank and R. P. Abelson.Scripts, Plans, Goals, and Understanding: An Inquiry Into Human Knowledge Structures. Psychology Press, New York, 1977. ISBN 9780203781036. doi: 10.4324/9780203781036
-
[54]
Sennheiser sk 20 bodypack transmitter xs wireless series,
Sennheiser. Sennheiser sk 20 bodypack transmitter xs wireless series, . URL https://docs. cloud.sennheiser.com/en-us/xsw/xsw/manual-sk-overview.html . Accessed: 2026- 04-22
2026
-
[55]
Sennheiser sk 2000 bodypack transmitter,
Sennheiser. Sennheiser sk 2000 bodypack transmitter, . URL https://www.sennheiser. com/en-us/catalog/products/wireless-systems/sk-2000. Accessed: 2026-04-22
2000
-
[56]
Setti, C
F. Setti, C. Russell, C. Bassetti, and M. Cristani. F-formation detection: Individuating free- standing conversational groups in images.PLoS ONE, 10, 2015
2015
-
[57]
Vista omni anchor,
Sewio. Vista omni anchor, . URL https://docs.sewio.net/docs/ anchor-vista-omni-30147663.html. Accessed: 2026-04-22
2026
-
[58]
Leonardo personal tag,
Sewio. Leonardo personal tag, . URL https://docs.sewio.net/docs/ tag-leonardo-personal-30146967.html. Accessed: 2026-04-22
2026
-
[59]
Swofford, J
M. Swofford, J. Peruzzi, N. Tsoi, S. Thompson, R. Martín-Martín, S. Savarese, and M. Vázquez. Improving social awareness through dante: Deep affinity network for clustering conversational interactants.Proceedings of the ACM on Human-Computer Interaction, 4(CSCW1):1–23, 2020
2020
-
[60]
S. Tan, D. M. Tax, and H. Hung. Conversation group detection with spatio-temporal con- text. InProceedings of the 2022 International Conference on Multimodal Interaction, ICMI ’22, page 170–180, New York, NY , USA, 2022. Association for Computing Machinery. ISBN 9781450393904. doi: 10.1145/3536221.3556611. URL https://doi.org/10.1145/ 3536221.3556611. 13
arXiv 2022
-
[61]
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295, 2024
Pith/arXiv arXiv 2024
-
[62]
Umagami, L
R. Umagami, L. Yue, X. Chu, R. Fukushima, T. Narita, Y . Mukuta, T. Takahata, J. Yang, and T. Harada. Intend to move: A multimodal dataset for intention-aware human motion understanding. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track
-
[63]
Umagami, L
R. Umagami, L. Yue, X. Chu, R. Fukushima, T. Narita, Y . Mukuta, T. Takahata, J. Yang, and T. Harada. Intend to move: A multimodal dataset for intention-aware human motion understanding. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2026. URL https://openreview.net/forum? id=3CVU3RRvPx
2026
-
[64]
J. Vargas. Covfee: Continuous Video Feedback Tool. URL https://github.com/josedvq/ covfee. Accessed: 2021-05-28
2021
-
[65]
Vargas Quiros, S
J. Vargas Quiros, S. Tan, C. Raman, E. Gedik, I. Pronotaris, and H. Hung. Spcl mingle badge. URL https://github.com/TUDelft-SPC-Lab/spcl_midge_hardware . Accessed: 2026- 04-22
2026
-
[66]
S. Vascon, E. Z. Mequanint, M. Cristani, H. Hung, M. Pelillo, and V . Murino. Detecting conversa- tional groups in images and sequences: A robust game-theoretic approach.Computer Vision and Image Understanding, 143:11–24, Feb. 2016. ISSN 1077-3142. doi: 10.1016/j.cviu.2015.09.012. URLhttps://www.sciencedirect.com/science/article/pii/S1077314215002076
-
[67]
B. Vissa. Agency in action: Entrepreneurs’ networking style and initiation of economic exchange.Organization Science, 23(2):492–510, 2012
2012
-
[68]
Y . Wu, J. Xiong, and X. Deng. How Social is It? A Benchmark for LLMs’ Capabilities in Multi- user Multi-turn Social Agent Tasks, Apr. 2025. URL http://arxiv.org/abs/2505.04628. arXiv:2505.04628 [cs]
Pith/arXiv arXiv 2025
-
[69]
Y . Xu, J. Zhang, Q. Zhang, and D. Tao. Vitpose: Simple vision transformer baselines for human pose estimation.Advances in neural information processing systems, 35:38571–38584, 2022
2022
-
[70]
H. Zhang, H. Xu, X. Wang, Q. Zhou, S. Zhao, and J. Teng. Mintrec: A new dataset for multimodal intent recognition. InProceedings of the 30th ACM International Conference on Multimedia, MM ’22, page 1688–1697. ACM, Oct. 2022. doi: 10.1145/3503161.3547906. URL http://dx.doi.org/10.1145/3503161.3547906
arXiv 2022
-
[71]
Zhang, X
H. Zhang, X. Wang, H. Xu, Q. Zhou, K. Gao, J. Su, jinyue Zhao, W. Li, and Y . Chen. MIntrec2.0: A large-scale benchmark dataset for multimodal intent recognition and out-of-scope detection in conversations. InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=nY9nITZQjc
2024
-
[72]
Zhang, Z
H. Zhang, Z. Li, Y . Zhu, H. Xu, P. Wang, H. Zhu, J. Zhou, and J. Zhang. Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark, Apr
-
[73]
S. Zhang, P. Sun, S. Chen, M. Xiao, W. Shao, W. Zhang, Y . Liu, K. Chen, and P. Luo. GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest. In A. Del Bue, C. Canton, J. Pont-Tuset, and T. Tommasi, editors,Computer Vision – ECCV 2024 Workshops, pages 52–70, Cham, 2025. Springer Nature Switzerland. ISBN 978-3-031-91813-1. doi: 10.1007/978-3...
-
[74]
X. Zhou, H. Zhu, L. Mathur, R. Zhang, H. Yu, Z. Qi, L.-P. Morency, Y . Bisk, D. Fried, G. Neubig, and M. Sap. SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents, Mar
-
[78]
Do you see multiple possibilities? Click ‘‘+’’to add more entries if you see multiple intentions or interpretations
-
[79]
Use your first impression and intuition; there is no right or wrong answer
Give your reasoning Complete the questionnaire under the video. Use your first impression and intuition; there is no right or wrong answer
-
[80]
If you have annotated the full video, you may submit
Continue the search Resume the video and repeat this process for every new intention you see until the clip ends. If you have annotated the full video, you may submit. **Grading rubric:** Your response will be evaluated based on the following criteria: Intention (not just actions): Describe what the participant is trying to achieve, not just what they are...
-
[81]
Adjust the timestamp to mark the exact start and end of the intention
Watch and pause Watch the clip and pause as soon as you notice an intention. Adjust the timestamp to mark the exact start and end of the intention. 24
-
[83]
No intention seen
Give your reasoning Complete the questionnaire under the video. Use your first impression and intuition; there is no right or wrong answer. Intention (Free text): Confidence: On a scale of 1-5, how confident are you in this interpretation? Just a guess / Extremely confident [Likert 1-5] Why? Provide the evidence from the video or audio that led you to thi...
-
[84]
Adjust the timestamp to mark the exact start and end of the intention
Watch and pause Watch the clip and pause as soon as you notice an intention. Adjust the timestamp to mark the exact start and end of the intention
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.