Pith. sign in

REVIEW 4 major objections 6 minor 93 references

This paper argues that perceived social intentions are inherently multiple and perspective-dependent, and presents COSI-Lab, a dataset that aligns participants' long-term goals with second-level, multi-observer 'apparent intention' annotati

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 00:48 UTC pith:C22Q3NW2

load-bearing objection A genuinely useful dataset resource for social-intention research, but its core AII annotations are unvalidated, so the headline coupling claim is a promise rather than a demonstrated result. the 4 major comments →

arxiv 2607.28649 v1 pith:C22Q3NW2 submitted 2026-06-02 cs.HC cs.AI

COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention

classification cs.HC cs.AI
keywords apparent intent inferencemulti-perspective annotationweakly scripted social interactionmultimodal datasetintention perceptionperspectivismconversation group detectiongoal hierarchy
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper introduces COSI-Lab, a multimodal dataset of a real academic networking workshop, and claims it is the first resource to couple participants' self-reported long-term goals with second-level annotations of how outside observers perceive their immediate intentions. The central argument is that perceived intentions are subjective and multiple, and should be modeled as explainable, perspective-driven reasoning rather than as label noise. The paper contributes an annotation protocol that elicits intention narratives together with the cues, assumptions, and beliefs that support them, from many annotators, plus benchmarks for intention narrative inference and conversation group detection. If the claim holds, the dataset enables a new line of research connecting goals, behavior, and perception of intent in ecologically valid settings.

Core claim

The paper argues that in weakly scripted social settings, a person's apparent intention is not a single ground-truth label but a space of plausible interpretations that different perceivers construct from observable cues, situation assumptions, and their own interpretative tendencies. COSI-Lab operationalizes this by having crowd-sourced annotators watch 30-second overhead video clips of real mingling and write free-form intention narratives with timestamps, confidence, intensity, and counterfactual alternatives, while also capturing annotator traits and demographics. The resulting dataset is claimed to be the first to align these multi-perspective apparent-intention annotations with the obs

What carries the argument

The Apparent Intent Inference (AII) problem: the task of inferring, from an ex-situ third-person perspective, what intention a person appears to have at a given moment, independent of whether that intention is later realized. The annotation protocol is the load-bearing instrument: it structures narratives through cues, situation characteristics, situation classes, and social scripts, and adds annotator-level trait data so that multiplicity of interpretations can be studied as a function of the perceiver rather than as noise.

Load-bearing premise

The claim depends on ex-situ online annotators' free-form narratives being a valid window into genuine variation in intention perception, rather than arbitrary responses produced by poorly motivated crowd workers.

What would settle it

Show that annotators' perceived intentions do not track any behavioral or self-report signal, for instance that participants' post-session goal-attainment ratings share no systematic association with the apparent intentions narrated by observers, or that two annotators' narratives for the same clip are no more similar to each other than to a shuffled baseline. More directly, if participants had reported their immediate proximal intentions during the event and observer narratives matched those reports at chance, the resource's validity as intention-perception data would collapse.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Researchers can study the relationship between self-reported long-term goals and apparent proximal intentions on the same individuals at the same event for the first time.
  • Intention-aware systems can be trained or evaluated to output multiple plausible intention hypotheses with explanations, rather than being forced to commit to a single label.
  • The dataset provides benchmark tasks for apparent intent inference and conversation group detection, including a human-LLM comparison that reveals systematic differences in how people and models describe intentions.
  • Multi-perspective annotations coupled with annotator traits allow investigation of how demographics and reflective functioning shape intention perception, supporting perspectivist modeling.
  • The weakly scripted, ecologically valid setting with real professional consequences makes findings more likely to transfer to in-the-wild social interactions than strongly scripted datasets.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the perspective-driven framing is correct, agreement between annotators should not be the primary quality metric; instead, downstream systems could be evaluated on the diversity and plausibility of the interpretations they generate, a departure from typical label-agreement benchmarks.
  • The dataset could be extended with participants' own in-the-moment proximal intention self-reports, for example via brief experience sampling between sessions, to directly validate whether ex-situ observer narratives track the observed person's actual immediate intentions.
  • The goal-to-perceived-intention link opens a route to study unrealized intentions: cases where observers perceive a goal-directed intention that the participant later reports not achieving, a key gap the paper identifies in existing intention research.
  • Because annotator traits and demographics are collected, the data could support deliberately resampling perspectives to make intention-inference systems aware of who is perceiving, potentially reducing demographic bias in social AI.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. COSI-Lab is a multimodal dataset of a two-session conference mingling event involving 32 academics, with overhead video, close-talk and wearable audio, IMU, UWB, surveys, and annotations. The paper's main contribution is an annotation protocol that elicits multi-perspective 'apparent intention' narratives from crowd annotators alongside self-reported participant goals, and the claim that this is the first resource coupling long-term goals with short-term third-party perceived intentions in a weakly scripted social setting. The paper also presents descriptive analyses of conversation groups, topics, and goals, plus benchmark demonstrations for AII and conversation-group detection.

Significance. If the AII annotations are meaningful, the dataset is a potentially valuable resource for perspectivist, explainable intention-inference research: it includes synchronized multimodal data, privacy-preserving processing, an annotation protocol with theory-grounded components (cues, assumptions, scripts), and an honest LLM-as-judge evaluation that reports a 45.3% discrimination accuracy. The detailed data collection and processing description, camera-wise 5-fold group detection evaluation, and released code for reproducibility are strengths. However, the core value rests on the validity of the AII narratives, and that validity is not yet demonstrated.

major comments (4)
  1. [§5.1, §7.1, §8 (Cue Grounding)] The load-bearing claim—that COSI-Lab couples self-reported long-term goals with multi-perspective apparent-intention annotations—is not supported by any reliability or validity evidence for the AII annotations. There is no inter-annotator agreement metric, no verification that cues cited in the narratives are present in the video/audio, and no analysis relating the AII narratives to the observed participants' own self-reported goals. Section 8 states that cue grounding 'was not assessed' and that goals and AII 'bare some relationship' without testing it. This makes the central dataset contribution, and desiderata D1–D4 built on it, currently an unverified assertion. At minimum, the paper should report inter-annotator agreement (e.g., Krippendorff's alpha on intention categories or annotation components), a cue-presence check on a sample, and an exploratory comparison between AII narrativ
  2. [§6, Appendix E.2/E.3] The quantitative descriptive claims in Figure 4 (topic composition and topic-redirection uptake) are derived entirely from LLM-assisted coding with no human validation. The appendix provides detailed prompts, but no inter-coder reliability against human coders, and no two-human agreement baseline. As reported, these statistics are trustworthy only insofar as the Qwen3-14B labels are reliable, which is not established. The authors should validate a sample of the LLM coding against human coders and report agreement; otherwise these results should be presented as illustrative rather than quantitative findings.
  3. [§7.1] The 'intention estimation' benchmark is currently descriptive, not predictive. It compares human and Gemma4-generated narratives via semantic similarity and an LLM-as-judge accuracy of 45.3%, but no concrete task formulation, evaluation metric, or baseline protocol is given that future models could be measured against. The 45.3% figure is reported without variance or confidence intervals and without a significance test, so its interpretation is unclear. To claim a benchmark task, the paper should define an evaluation setup (e.g., predicting annotation components, matching a target perspective, or ranking narrative plausibility) and provide a reproducible metric.
  4. [§1, Table 1, §9] The claim of satisfying desideratum D4 ('hierarchical structure of human intention') and the conclusion that COSI-Lab 'for the first time' enables study of 'subgoals or apparent proximal intentions' go beyond what is demonstrated. The dataset contains self-reported long-term goals and AII for short clips, but no annotation explicitly linking proximal intentions to subgoals or to distal goal hierarchies. The checkmark for D4 in Table 1 and the conclusion wording should be softened or accompanied by an explicit description of which hierarchy-related variables are present in the released data.
minor comments (6)
  1. [Abstract] 'a explainable' should be 'an explainable'.
  2. [§5.1] 'the begging and end of the intention' should be 'the beginning and end'.
  3. [§8] 'they bare some relationship' should be 'they bear some relationship'.
  4. [Appendix C.2.1] 'felxible inputs' should be 'flexible inputs'.
  5. [Appendix B] 'the michrophones record' should be 'the microphones record'.
  6. [§7.1 / Figure 5] The semantic-similarity plots are described qualitatively as showing 'clear clustering separation'; a quantitative cluster-separation or distance statistic would make the claim more precise.

Circularity Check

0 steps flagged

No significant circularity: COSI-Lab is a resource paper whose benchmarks and analyses are self-contained or explicitly deferred.

full rationale

The paper's central contribution is the COSI-Lab dataset itself—a multimodal recording of a weakly scripted workshop with self-reported goals, multi-perspective AII annotations, and benchmark tasks. There is no derivation chain in which a predicted quantity is constructed from the same quantity it claims to predict, and no fitted parameter is renamed as a prediction. The AII annotations are elicited from crowd annotators and then compared with LLM outputs; the LLM-as-judge result (45.3%) is reported as a failure to distinguish human from model narratives, not as evidence for model superiority. The conversation-group detection benchmark trains standard models (DANTE, LSTM) on position/orientation features and evaluates against human group annotations; this is a conventional supervised evaluation, not circular. The descriptive topic-redirection analyses do rely on LLM-generated labels without human validation, which is a validity concern, but it is not circularity: the LLM is not being used to 'predict' a label that was itself derived from the LLM in a way that defines the result. The paper also openly concedes limitations that would be relevant to correctness or validity but not to circularity: cue grounding is not assessed, and the relationship between self-reported goals and AII is asserted rather than tested. Self-citations to ConfLab [46] and related prior work place the dataset in a lineage but do not function as an unverified load-bearing premise for the paper's conclusions. Overall, no step reduces to its own inputs by construction, so the circularity score is 0.

Axiom & Free-Parameter Ledger

4 free parameters · 6 axioms · 0 invented entities

The central contribution is the dataset itself, so free parameters are mostly choices in the processing/benchmark pipeline rather than fitted constants in a derivation. The AII task framing is a problem definition, not an invented entity; no independent evidence is needed. The main burden is the validity of subjective annotations and the unvalidated LLM coding.

free parameters (4)
  • Assumed body height H = 1.7 m
    Used in Appendix C.2.3 for ray-plane intersection to reconstruct 3D keypoints; hand-chosen constant applied to all participants.
  • Keypoint height ratios α_i = from COCO model (not given)
    Used with H to assign a world height to each keypoint in triangulation; standard model-derived ratios, not fitted to this dataset.
  • LLM coding window parameters = 300-turn window, 250-turn stride (topic distribution); 300/220/80 for redirection screening
    Hand-chosen parameters for the LLM-assisted descriptive coding of topics and redirection uptake; different choices could affect reported percentages.
  • Frame strides for baselines = 20 (LSTM), 300 (DANTE)
    Downsampling rates for the two group-detection baselines; chosen by hand and affect benchmark runtime and accuracy.
axioms (6)
  • domain assumption The workshop setting is ecologically valid for studying social intentions despite being recorded and weakly scripted.
    Used to argue that captured interactions reflect real professional networking (Sec. 3.1).
  • domain assumption Head orientation is a valid indicator of conversational attention and group membership (Kendon's F-formation).
    Basis for conversation-group annotation in Sec. 5.2 and for using head position/orientation in benchmarks (F.2).
  • domain assumption 30-second clips are sufficient for a third-party observer to perceive proximal intentions.
    Annotation design cuts clips to 30s (Sec. 5.1, F.1); if too short, AII narratives would be impoverished.
  • ad hoc to paper LLM-generated topic/redirection labels are acceptably reliable for descriptive statistics without human-validation metrics.
    Appendix E.2/E.3 uses Qwen3-14B for topic distribution and redirection uptake but reports no agreement with human coders; the paper calls outputs 'validated' only by another LLM.
  • domain assumption If multiple plausible intention narratives are socially meaningful, evaluating them requires diversity/grounding/plausibility measures.
    The narrative evaluation framework in Sec. 5.1 and 7.1 assumes these three dimensions are the right ones.
  • standard math Hyperbolic Tangent Similarity (HTS) is a valid semantic similarity metric.
    Used to quantify diversity of annotations; relies on [43] and an embedding model.

pith-pipeline@v1.3.0-alltime-deepseek · 29734 in / 13236 out tokens · 115078 ms · 2026-08-03T00:48:15.757039+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention." pith.science (2026). https://pith.science/paper/C22Q3NW2

@misc{pith2026260728649,
  author       = {Pith},
  title        = {Pith review of: COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/C22Q3NW2}},
  note         = {Machine review of arXiv:2607.28649}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

COSI-Lab presents a multimodal, multi-sensor dataset of an interdisciplinary scientific workshop containing 32 academics at an international conference. It captures ecologically valid social interactions in a weakly scripted setting consisting of two 30-minute mingling sessions with real professional and social consequences for the participants involved. We argue that future intelligent systems could be better equipped to handle subjective perceptions by modeling their multiplicity not as label noise but as a explainable perspective-driven reasoning process. We focus on the Apparent Intent Inference (AII) problem as determined by ex-situ observers and conceptualize intentions to be independent of manifest future outcomes. We contribute 1. a novel annotation process for AII that accounts for a perceiver's own interpretative tendencies, 2. quantitative and qualitative analyses of intent narratives with respect to diversity, grounding, and plausibility; 3. benchmark tasks for AII and surrounding relevant contextual factors such as social involvement; 4. speech quality audio for all participants as well as privacy preserving multi-modal data, enabling lexical and nonverbal behavior analysis; and 5. coupling of self-reported goals of each participant (30 minute to 3 hour) with annotated AII (seconds).

Figures

Figures reproduced from arXiv: 2607.28649 by Anne L.J. ter Wal, Arthur Mercier, Balint Dioszegi, Bernd Dudzik, Chenxu Hao, Chirag Raman, Gara Dorta, Hayley Hung, Ivan Kondyurin, Jorge Castro-God\'inez, Jose Morales-Vargas, Laura Cabrera-Quir\'os, Litian Li, Nale Lehmann-Willenbrock, Saunaq Chakrabarty, Sotiris Vacanas, Stephanie Tan, Vanessa Begemann, Vitaliy Popov, Zonghuan Li.

Figure 1
Figure 1. Figure 1: COSI-Lab: The conference living lab for studying multi-perspective social intentions. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Workshop procedure as predefined tasks where the structure of activity is specified and strongly-scripted in nature, which limits the possible emergent intentions that could arise in more open contexts. Multimodal human behavior Existing multimodal human behavior datasets [46, 11, 1] are typ￾ically not focused on social intentions, and also lacking either ecological validity or modalities. MIntRec2.0 focus… view at source ↗
Figure 3
Figure 3. Figure 3: Interaction structure and survey-reported goals across the two mingling sessions. [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Conversation topic composition and redirection uptake by session and group size. M1/M2 [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Left: Semantic similarity of the annotations and presence of emotions, cues, beliefs, [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Dataset file structure. For more information, refer to the corresponding [PITH_FULL_IMAGE:figures/full_fig_p016_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Example calibration footage for camera geometry estimation. The left panel shows a per [PITH_FULL_IMAGE:figures/full_fig_p017_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Illustration of the process of benchmarking task 2. (a)Segmentation mask from SAM3 [PITH_FULL_IMAGE:figures/full_fig_p020_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Displayed a screen grab of the UI of Covfee. [PITH_FULL_IMAGE:figures/full_fig_p022_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

93 extracted references · 5 canonical work pages

  1. [1]

    Agrawal, A

    V . Agrawal, A. Akinyemi, K. Alvero, M. Behrooz, J. Buffalini, F. M. Carlucci, J. Chen, J. Chen, Z. Chen, S. Cheng, et al. Seamless interaction: Dyadic audiovisual motion modeling and large-scale dataset.arXiv preprint arXiv:2506.22554, 2025

  2. [2]

    Akata, D

    Z. Akata, D. Balliet, M. De Rijke, F. Dignum, V . Dignum, G. Eiben, A. Fokkens, D. Grossi, K. Hindriks, H. Hoos, et al. A research agenda for hybrid intelligence: augmenting human intel- lect with collaborative, adaptive, responsible, and explainable artificial intelligence.Computer, 53(8):18–28, 2020

  3. [3]

    Alameda-Pineda, J

    X. Alameda-Pineda, J. Staiano, R. Subramanian, L. M. Batrinca, E. Ricci, B. Lepri, O. Lanz, and N. Sebe. SALSA: A Novel Dataset for Multimodal Group Behavior Analysis.CoRR, abs/1506.06882, 2015. URLhttp://arxiv.org/abs/1506.06882

  4. [4]

    M. C. Ashton and K. Lee. The hexaco–60: A short measure of the major dimensions of personality.Journal of personality assessment, 91(4):340–345, 2009

  5. [5]

    Z. Bai, P. Wang, T. Xiao, T. He, Z. Han, Z. Zhang, and M. Z. Shou. Hallucination of multimodal large language models: A survey.arXiv preprint arXiv:2404.18930, 2024

  6. [6]

    M. Bain, J. Huh, T. Han, and A. Zisserman. Whisperx: Time-accurate speech transcription of long-form audio.arXiv preprint arXiv:2303.00747, 2023

  7. [7]

    Belardinelli

    A. Belardinelli. Gaze-based intention estimation: principles, methodologies, and applications in hri, 2023

  8. [8]

    Bratman.Intention, plans, and practical reason

    M. Bratman.Intention, plans, and practical reason. Harvard University Press, Cambridge, MA, 1987

  9. [9]

    Broekens and W.-P

    J. Broekens and W.-P. Brinkman. Affectbutton: A method for reliable and valid affective self-report.International Journal of Human-Computer Studies, 71(6):641–667, 2013

  10. [10]

    Cabitza, A

    F. Cabitza, A. Campagner, and V . Basile. Toward a perspectivist turn in ground truthing for predictive computing.Proceedings of the AAAI Conference on Artificial Intelligence, 37:6860–6868, 6 2023. ISSN 2374-3468. doi: 10.1609/AAAI.V37I6.25840. URL https: //ojs.aaai.org/index.php/AAAI/article/view/25840

  11. [11]

    Cabrera-Quiros, A

    L. Cabrera-Quiros, A. Demetriou, E. Gedik, L. van der Meij, and H. Hung. The matchnmingle dataset: a novel multi-sensor resource for the analysis of social interactions and group dynamics in-the-wild during free-standing conversations and speed dates.IEEE Transactions on Affective Computing, 12(1):113–130, 2018. 10

  12. [12]

    Cabrera-Quiros, A

    L. Cabrera-Quiros, A. Demetriou, E. Gedik, L. van der Meij, and H. Hung. The matchnmingle dataset: A novel multi-sensor resource for the analysis of social interactions and group dynamics in-the-wild during free-standing conversations and speed dates.IEEE Transactions on Affective Computing, 12(1):113–130, 2021

  13. [13]

    Calib.io calibrator

    calib.io. Calib.io calibrator. URL https://calib.io/products/calib. Accessed: 2026- 04-22

  14. [14]

    Carion, L

    N. Carion, L. Gustafson, Y .-T. Hu, S. Debnath, R. Hu, D. Suris, C. Ryali, K. V . Alwala, H. Khedr, A. Huang, et al. Sam 3: Segment anything with concepts.arXiv preprint arXiv:2511.16719, 2025

  15. [15]

    J. M. Cheek and A. H. Buss. Shyness and sociability.Journal of personality and social psychology, 41(2):330, 1981

  16. [16]

    K. Chen, Z. Zhang, W. Zeng, R. Zhang, F. Zhu, and R. Zhao. Shikra: Unleashing Multimodal LLM’s Referential Dialogue Magic, July 2023. URL http://arxiv.org/abs/2306.15195. arXiv:2306.15195 [cs]

  17. [17]

    K. M. Connor, J. R. Davidson, L. E. Churchill, A. Sherwood, R. H. Weisler, and E. Foa. Psychometric properties of the social phobia inventory (spin): New self-rating scale.The British Journal of Psychiatry, 176(4):379–386, 2000

  18. [18]

    Dudzik and J

    B. Dudzik and J. Broekens. A valid self-report is never late, nor is it early: On considering the "right" temporal distance for assessing emotional experience, 2023

  19. [19]

    Dudzik, M.-P

    B. Dudzik, M.-P. Jansen, F. Burger, F. Kaptein, J. Broekens, D. K. Heylen, H. Hung, M. A. Neerincx, and K. P. Truong. Context in human emotion perception for automatic affect detection: A survey of audiovisual databases. In2019 8th International Conference on Affective Computing and Intelligent Interaction (ACII), pages 206–212, Cambridge, UK, 2019. IEEE

  20. [20]

    C. Edelsky. Who’s Got the Floor?Language in Society, 10(3):383–421, 1981. ISSN 0047-4045

  21. [21]

    F. F, K. I, S. K, and W. C. Toward a script theory of guidance in computer-supported collaborative learning.Educ Psychol, 2013. doi: 10.1080/00461520.2012.748005

  22. [22]

    Fonagy, P

    P. Fonagy, P. Luyten, A. Moulton-Perkins, Y .-W. Lee, F. Warren, S. Howard, R. Ghinai, P. Fearon, and B. Lowyck. Development and validation of a self-report measure of mentalizing: The reflective functioning questionnaire.PloS one, 11(7):e0158678, 2016

  23. [23]

    Gebru, J

    T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. D. Iii, and K. Crawford. Datasheets for datasets.Communications of the ACM, 64(12):86–92, 2021

  24. [24]

    Grauman, A

    K. Grauman, A. Westbury, L. Torresani, K. Kitani, J. Malik, T. Afouras, K. Ashutosh, V . Baiyya, S. Bansal, B. Boote, et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19383–19400. IEEE Computer Society, 2024

  25. [25]

    J. Gu, X. Jiang, Z. Shi, H. Tan, X. Zhai, C. Xu, W. Li, Y . Shen, S. Ma, H. Liu, S. Wang, K. Zhang, Y . Wang, W. Gao, L. Ni, and J. Guo. A survey on llm-as-a-judge, 2025. URL https://arxiv.org/abs/2411.15594

  26. [26]

    Honnibal, I

    M. Honnibal, I. Montani, S. Van Landeghem, A. Boyd, et al. spacy: Industrial-strength natural language processing in python. 2020

  27. [27]

    T. M. Hrkalovic, B. Dudzik, D. Balliet, and H. Hung. Parsel: A multimodal dataset for modeling decision-making processes involved in selecting partners for joint tasks.IEEE Transactions on Affective Computing, 2025

  28. [28]

    Hung and B

    H. Hung and B. Kröse. Detecting F-formations as dominant sets. InProceedings of the 13th international conference on multimodal interfaces, pages 231–238, Alicante Spain, Nov. 2011. ACM. ISBN 978-1-4503-0641-6. doi: 10.1145/2070481.2070525. URL https://dl.acm. org/doi/10.1145/2070481.2070525. 11

  29. [29]

    H. Hung, L. Li, J. Molhoek, and J. Zhou. The discontent with intent estimation in-the-wild: the case for unrealized intentions. InExtended abstracts of the CHI conference on human factors in computing systems, pages 1–9, 2024

  30. [30]

    life of the party

    P. Ingram and M. W. Morris. Do people mix at mixers? structure, homophily, and the “life of the party”.Administrative Science Quarterly, 52(4):558–585, 2007

  31. [31]

    T. Jing, T. Chen, R. Tian, Y . Chen, J. Domeyer, H. Toyoda, R. Sherony, and Z. Ding. Psi: A benchmark for human interpretation and response in traffic interactions. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2025

  32. [32]

    Kendon.Conducting Interaction: Patterns of Behavior in Focused Encounters

    A. Kendon.Conducting Interaction: Patterns of Behavior in Focused Encounters. Cambridge University Press, Cambridge, UK, 1990

  33. [33]

    Khindkar, V

    V . Khindkar, V . Balasubramanian, C. Arora, A. Subramanian, and C. Jawahar. Can reasons help improve pedestrian intent estimation? a cross-modal approach. In2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 11515–11522. IEEE, 2024

  34. [34]

    A. W. Kruglanski, M. Chernikova, M. Babush, M. Dugas, and B. M. Schumpe. Chapter three - the architecture of goal systems: Multifinality, equifinality, and counterfinality in means—end relations. volume 2 ofAdvances in Motivation Science, pages 69–98. Elsevier, 2015. doi: 10.1016/bs.adms.2015.04.001

  35. [35]

    Kuchaiev, J

    O. Kuchaiev, J. Li, H. Nguyen, O. Hrinchuk, R. Leary, B. Ginsburg, S. Kriman, S. Beliaev, V . Lavrukhin, J. Cook, P. Castonguay, M. Popova, J. Huang, and J. M. Cohen. Nemo: a toolkit for building ai applications using neural modules, 2019

  36. [36]

    R. D. Lennox and R. N. Wolfe. Revision of the self-monitoring scale. 1984

  37. [37]

    J. Li, P. Wei, W. Han, and L. Fan. Intentqa: Context-aware video intent reasoning. InProceedings of the IEEE/CVF international conference on computer vision, pages 11963–11974, 2023

  38. [38]

    Z. Lian, H. Chen, L. Chen, H. Sun, L. Sun, Y . Ren, Z. Cheng, B. Liu, R. Liu, X. Peng, J. Yi, and J. Tao. AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models, May 2025. URL http://arxiv.org/abs/2501. 16566. arXiv:2501.16566 [cs]

  39. [39]

    Martín-Fernández, B

    M. Martín-Fernández, B. Requero, X. Zhou, D. Gonçalves, and D. Santos. Refinement of the analysis-holism scale: A cross-cultural adaptation and validation of two shortened measures of analytic versus holistic thinking in spain and the united states.Personality and Individual Dif- ferences, 186:111322, 2022. ISSN 0191-8869. doi: https://doi.org/10.1016/j.p...

  40. [40]

    Mathur, M

    L. Mathur, M. Qian, P. P. Liang, and L.-P. Morency. Social Genome: Grounded Social Reasoning Abilities of Multimodal Models, Feb. 2025. URL http://arxiv.org/abs/2502.15109. arXiv:2502.15109 [cs]

  41. [41]

    A. R. Mele.Springs of action: Understanding intentional behavior. Oxford University Press, New York, USA, 1992

  42. [42]

    Presidio: Data Protection and De-identification SDK

    Microsoft. Presidio: Data Protection and De-identification SDK. https://microsoft. github.io/presidio/

  43. [43]

    V . S. R. Parupudi. Magnitude matters: a superior class of similarity metrics for holistic semantic understanding, 2025. URLhttps://arxiv.org/abs/2509.19323

  44. [44]

    Z. Peng, W. Wang, L. Dong, Y . Hao, S. Huang, S. Ma, and F. Wei. Kosmos-2: Grounding Multimodal Large Language Models to the World, July 2023. URL http://arxiv.org/abs/ 2306.14824. arXiv:2306.14824 [cs]

  45. [45]

    J. V . Quiros, C. Raman, S. Tan, E. Gedik, L. Cabrera-Quiros, and H. Hung. Rewind dataset: Privacy-preserving speaking status segmentation from multimodal body movement signals in the wild, 2024. 12

  46. [46]

    Raman, J

    C. Raman, J. Vargas Quiros, S. Tan, A. Islam, E. Gedik, and H. Hung. Conflab: A data collection concept, dataset, and benchmark for machine analysis of free-standing social interactions in the wild. In S. Koyejo, S. Mohamed, A. Agarwal, D. Bel- grave, K. Cho, and A. Oh, editors,Advances in Neural Information Processing Sys- tems, volume 35, pages 23701–23...

  47. [47]

    Rasouli, I

    A. Rasouli, I. Kotseruba, T. Kunic, and J. K. Tsotsos. Pie: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction. InInternational Conference on Computer Vision (ICCV), pages 6262–6271, Seoul, South Korea, 2019. IEEE

  48. [48]

    J. F. Rauthmann and R. A. Sherman. Chapter 13 - conceptualizing and measuring the psychological situation. In D. Wood, S. J. Read, P. Harms, and A. Slaughter, editors, Measuring and Modeling Persons and Situations, pages 427–463. Academic Press, 2021. ISBN 978-0-12-819200-9. doi: https://doi.org/10.1016/B978-0-12-819200-9.00009-0. URL https://www.scienced...

  49. [49]

    Rme 32 ad,

    RME. Rme 32 ad, . URL https://rme-audio.de/m-32-m-16-ad.html . Accessed: 2026- 04-22

  50. [50]

    Rme fireface ufx iii,

    RME. Rme fireface ufx iii, . URL https://rme-audio.de/fireface-ufx-3.html . Ac- cessed: 2026-04-22

  51. [51]

    R. B. Rubin and M. M. Martin. Development of a measure of interpersonal communication competence.Communication Research Reports, 11(1):33–44, 1994

  52. [52]

    Rudenko, L

    A. Rudenko, L. Palmieri, M. Herman, K. M. Kitani, D. M. Gavrila, and K. O. Arras. Human motion trajectory prediction: a survey.The International Journal of Robotics Research, 39 (8):895–935, 2020. doi: 10.1177/0278364920917446. URL https://doi.org/10.1177/ 0278364920917446

  53. [53]

    R. C. Schank and R. P. Abelson.Scripts, Plans, Goals, and Understanding: An Inquiry Into Human Knowledge Structures. Psychology Press, New York, 1977. ISBN 9780203781036. doi: 10.4324/9780203781036

  54. [54]

    Sennheiser sk 20 bodypack transmitter xs wireless series,

    Sennheiser. Sennheiser sk 20 bodypack transmitter xs wireless series, . URL https://docs. cloud.sennheiser.com/en-us/xsw/xsw/manual-sk-overview.html . Accessed: 2026- 04-22

  55. [55]

    Sennheiser sk 2000 bodypack transmitter,

    Sennheiser. Sennheiser sk 2000 bodypack transmitter, . URL https://www.sennheiser. com/en-us/catalog/products/wireless-systems/sk-2000. Accessed: 2026-04-22

  56. [56]

    Setti, C

    F. Setti, C. Russell, C. Bassetti, and M. Cristani. F-formation detection: Individuating free- standing conversational groups in images.PLoS ONE, 10, 2015

  57. [57]

    Vista omni anchor,

    Sewio. Vista omni anchor, . URL https://docs.sewio.net/docs/ anchor-vista-omni-30147663.html. Accessed: 2026-04-22

  58. [58]

    Leonardo personal tag,

    Sewio. Leonardo personal tag, . URL https://docs.sewio.net/docs/ tag-leonardo-personal-30146967.html. Accessed: 2026-04-22

  59. [59]

    Swofford, J

    M. Swofford, J. Peruzzi, N. Tsoi, S. Thompson, R. Martín-Martín, S. Savarese, and M. Vázquez. Improving social awareness through dante: Deep affinity network for clustering conversational interactants.Proceedings of the ACM on Human-Computer Interaction, 4(CSCW1):1–23, 2020

  60. [60]

    S. Tan, D. M. Tax, and H. Hung. Conversation group detection with spatio-temporal con- text. InProceedings of the 2022 International Conference on Multimodal Interaction, ICMI ’22, page 170–180, New York, NY , USA, 2022. Association for Computing Machinery. ISBN 9781450393904. doi: 10.1145/3536221.3556611. URL https://doi.org/10.1145/ 3536221.3556611. 13

  61. [61]

    G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295, 2024

  62. [62]

    Umagami, L

    R. Umagami, L. Yue, X. Chu, R. Fukushima, T. Narita, Y . Mukuta, T. Takahata, J. Yang, and T. Harada. Intend to move: A multimodal dataset for intention-aware human motion understanding. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track

  63. [63]

    Umagami, L

    R. Umagami, L. Yue, X. Chu, R. Fukushima, T. Narita, Y . Mukuta, T. Takahata, J. Yang, and T. Harada. Intend to move: A multimodal dataset for intention-aware human motion understanding. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2026. URL https://openreview.net/forum? id=3CVU3RRvPx

  64. [64]

    J. Vargas. Covfee: Continuous Video Feedback Tool. URL https://github.com/josedvq/ covfee. Accessed: 2021-05-28

  65. [65]

    Vargas Quiros, S

    J. Vargas Quiros, S. Tan, C. Raman, E. Gedik, I. Pronotaris, and H. Hung. Spcl mingle badge. URL https://github.com/TUDelft-SPC-Lab/spcl_midge_hardware . Accessed: 2026- 04-22

  66. [66]

    Vascon, E

    S. Vascon, E. Z. Mequanint, M. Cristani, H. Hung, M. Pelillo, and V . Murino. Detecting conversa- tional groups in images and sequences: A robust game-theoretic approach.Computer Vision and Image Understanding, 143:11–24, Feb. 2016. ISSN 1077-3142. doi: 10.1016/j.cviu.2015.09.012. URLhttps://www.sciencedirect.com/science/article/pii/S1077314215002076

  67. [67]

    B. Vissa. Agency in action: Entrepreneurs’ networking style and initiation of economic exchange.Organization Science, 23(2):492–510, 2012

  68. [68]

    Y . Wu, J. Xiong, and X. Deng. How Social is It? A Benchmark for LLMs’ Capabilities in Multi- user Multi-turn Social Agent Tasks, Apr. 2025. URL http://arxiv.org/abs/2505.04628. arXiv:2505.04628 [cs]

  69. [69]

    Y . Xu, J. Zhang, Q. Zhang, and D. Tao. Vitpose: Simple vision transformer baselines for human pose estimation.Advances in neural information processing systems, 35:38571–38584, 2022

  70. [70]

    Zhang, H

    H. Zhang, H. Xu, X. Wang, Q. Zhou, S. Zhao, and J. Teng. Mintrec: A new dataset for multimodal intent recognition. InProceedings of the 30th ACM International Conference on Multimedia, MM ’22, page 1688–1697. ACM, Oct. 2022. doi: 10.1145/3503161.3547906. URL http://dx.doi.org/10.1145/3503161.3547906

  71. [71]

    Zhang, X

    H. Zhang, X. Wang, H. Xu, Q. Zhou, K. Gao, J. Su, jinyue Zhao, W. Li, and Y . Chen. MIntrec2.0: A large-scale benchmark dataset for multimodal intent recognition and out-of-scope detection in conversations. InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=nY9nITZQjc

  72. [72]

    Zhang, Z

    H. Zhang, Z. Li, Y . Zhu, H. Xu, P. Wang, H. Zhu, J. Zhou, and J. Zhang. Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark, Apr

  73. [73]

    Zhang, P

    S. Zhang, P. Sun, S. Chen, M. Xiao, W. Shao, W. Zhang, Y . Liu, K. Chen, and P. Luo. GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest. In A. Del Bue, C. Canton, J. Pont-Tuset, and T. Tommasi, editors,Computer Vision – ECCV 2024 Workshops, pages 52–70, Cham, 2025. Springer Nature Switzerland. ISBN 978-3-031-91813-1. doi: 10.1007/978-3...

  74. [74]

    X. Zhou, H. Zhu, L. Mathur, R. Zhang, H. Yu, Z. Qi, L.-P. Morency, Y . Bisk, D. Fried, G. Neubig, and M. Sap. SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents, Mar

  75. [78]

    Do you see multiple possibilities? Click ‘‘+’’to add more entries if you see multiple intentions or interpretations

  76. [79]

    Use your first impression and intuition; there is no right or wrong answer

    Give your reasoning Complete the questionnaire under the video. Use your first impression and intuition; there is no right or wrong answer

  77. [80]

    If you have annotated the full video, you may submit

    Continue the search Resume the video and repeat this process for every new intention you see until the clip ends. If you have annotated the full video, you may submit. **Grading rubric:** Your response will be evaluated based on the following criteria: Intention (not just actions): Describe what the participant is trying to achieve, not just what they are...

  78. [81]

    Adjust the timestamp to mark the exact start and end of the intention

    Watch and pause Watch the clip and pause as soon as you notice an intention. Adjust the timestamp to mark the exact start and end of the intention. 24

  79. [83]

    No intention seen

    Give your reasoning Complete the questionnaire under the video. Use your first impression and intuition; there is no right or wrong answer. Intention (Free text): Confidence: On a scale of 1-5, how confident are you in this interpretation? Just a guess / Extremely confident [Likert 1-5] Why? Provide the evidence from the video or audio that led you to thi...

  80. [84]

    Adjust the timestamp to mark the exact start and end of the intention

    Watch and pause Watch the clip and pause as soon as you notice an intention. Adjust the timestamp to mark the exact start and end of the intention

Showing first 80 references.