REVIEW 4 major objections 6 minor 3 references
This paper introduces InterAct, a multimodal dataset of 241 minute-or-longer improvised scenes in which two actors perform realistic, emotionally labelled interactions while their audio, body motion, and facial expressions are captured simu
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-05 05:04 UTC pith:5B4FDJ6B
load-bearing objection InterAct is a genuinely useful dataset—first long-form dyadic dynamic interaction corpus with body, face, audio, and labels—but the paper overstates its uniqueness and the acting-realism gap is real, though honestly conceded. the 4 major comments →
InterAct: A Large-Scale Dataset of Dynamic, Expressive and Interactive Activities between Two People in Daily Scenarios
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that daily two-person interaction—not just conversation—can be captured as a single synchronized multimodal record, and that this changes what speech-driven generation can attempt. InterAct is built by giving pairs of actors a role setup and an emotion label and letting them improvise a complete task or scene for a minute or longer in a 5m x 5m capture volume; 241 such sequences were collected from 7 actors. Each sequence ships separated speech for both actors, 53 body markers plus 20 finger markers per actor, facial mesh sequences converted to ARKit blendshapes, frame-level action labels, and relationship and emotion metadata. Statistical analysis shows wider spreads of
What carries the argument
The load-bearing object is the InterAct capture pipeline and its annotation scheme: 28 VICON cameras tracking 53 body and 20 finger markers per actor, head-mounted iPhones recording face RGB-D and speech, wireless timecode synchronization between the motion-capture system and the phones, and a speech-separation step that yields clean per-actor audio. The scenario prompts, which pair a relationship with an emotion, are what make the data objective-driven and semantically coherent rather than free-form chatting. The baseline's two mechanisms—hierarchical diffusion that first regresses lower-body joints and then upper-body joints conditioned on them, and denoiser-weight swapping during face fin
Load-bearing premise
The dataset's value assumes that one-minute improvised performances by seven actors, prompted with roles and emotions, are representative enough of real daily two-person interactions that models trained on them carry over to natural behavior.
What would settle it
Train the released baseline on InterAct and evaluate it on unscripted natural dyads captured with ordinary video and pose estimation; if the generated relative-distance and orientation dynamics match those of static-conversation baselines rather than the natural dyads, or if human raters cannot distinguish InterAct-trained outputs from outputs trained on static conversational data, the acting-data representativeness claim is falsified.
If this is right
- Audio-driven generation can now target two people who walk, sit, shift position, and orient toward or away from each other, not just speakers rooted in place.
- The dataset supplies a benchmark in which models must reproduce long-range interaction patterns, proxemics, and relative body orientation, not only local gestures.
- The hierarchical body diffusion recipe suggests a practical way to split high-dimensional two-person motion generation into lower-body control and upper-body detail.
- The denoiser-swap fine-tuning trick gives a cheap way to improve lip sync on a small high-quality face dataset without retraining the whole model.
- Because each sequence carries emotion, relationship, and action labels, generated motion can be steered by those labels at inference time.
Where Pith is reading between the lines
- If improvisation under role prompts generalizes, the same capture protocol could be used with non-professional participants or scripted naturalistic tasks; a useful check would be comparing InterAct's relative-distance and orientation distributions against those of unscripted dyadic video.
- The lip-accuracy denoiser swap is separable from the rest of the face model in spirit; testing it inside a different speech-to-face architecture would show whether the trick transfers.
- The captured hand-contact statistics, such as higher hand-hand self-contact rates in friends and social scenarios, suggest InterAct could serve as a source for proxemics and nonverbal-communication studies, though the paper only reports the statistics and does not develop that analysis.
- The flat-ground, prop-free setup means models trained on the dataset may not capture object-mediated interactions; augmenting with object interaction data or synthetic props would be a natural test of transfer.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces InterAct, a two-person multimodal motion-capture dataset: 241 sequences (about 10 hours) of two actors improvising role- and emotion-prompted daily-life scenarios, with synchronized audio, body MoCap (53 markers plus finger markers), and facial ARKit blendshapes for both participants. The authors also propose a speech-conditioned diffusion baseline that hierarchically generates two people's body motions and facial expressions, and they evaluate it on held-out InterAct sequences with quantitative metrics and user studies. The dataset and code are promised for public release.
Significance. If the claims are corrected and the data are released, InterAct would be a useful resource for dyadic interaction research: it is one of the few datasets with simultaneous body, face, and audio capture of two interacting people in non-stationary scenarios, and it includes long (1+ minute) coherent activities. The capture pipeline is described in detail, statistical comparisons are provided, and the baseline is a reproducible attempt at a genuinely hard joint generation task. I agree with the paper that evaluating the baseline on held-out sequences from the same dataset is standard practice and not circular; however, the paper's novelty claims are currently overstated, and several external-validity claims go beyond what the evidence supports.
major comments (4)
- [Sec. 3.2 and Table 1] The claim that InterAct is 'the only one that captures both actors' body motions and facial expressions at the same time' is directly contradicted by the paper's own Table 1: the Ng et al. (2024) row lists Audio ✓, Hands Detailed, Face ✓, and Coherence ✓. Since the related-work section also describes Ng et al. as a two-person conversation dataset, the uniqueness claim must be narrowed to the specific combination of dynamic, large-range, two-person interaction with simultaneous face capture, or the comparison table must be corrected. This is load-bearing because the abstract and introduction frame InterAct's contribution as a first-of-kind resource.
- [Sec. 3.2 and Fig. 4 caption] The main text says Ng et al. (2024) 'records both actors primarily during conversation standing still,' but the Fig. 4 caption states that 'only one actor is captured in BEAT and in [Ng et al. 2024], so analysis in the lower part is not applicable on them.' These statements are mutually inconsistent and affect the central statistical comparison: whether InterAct uniquely enables two-person relative-position/orientation analysis depends on whether Ng et al. actually captured both participants' body motion. The authors must resolve this contradiction and, if both actors are available in Ng et al., include the lower-panel comparison or justify its exclusion.
- [Sec. 7.1 and Sec. 8] The dataset's value for 'human behavior grounding' and 'real-life, physically-informed data' (Sec. 8) rests on the assumption that improvised performances by 7 actors on a flat 5m x 5m stage, without props and sometimes with adults playing children, are representative of natural dyadic behavior. Section 7.1 concedes the acting nature but no validation is offered: the statistical analyses in Sec. 3.2 measure internal diversity, not fidelity to real interaction. I do not consider the use of acted data fatal by itself, since similar resources (BEAT, Ng et al.) are also acted or controlled, but the conclusion's wording overstates ecological validity. The authors should either provide an external comparison (e.g., proxemics or contact statistics against a naturalistic corpus such as TalkHands) or substantially soften the 'real-life' and 'socially-aware robotics' claims.
- [Sec. 5] All baseline evaluations, including Tables 2 and 4 and both user studies, are performed on a held-out portion of InterAct itself, i.e., in-distribution with the same actors and the same improvised acting style. The paper should state this limitation explicitly when claiming the method is 'effective' and avoid implying that generated motions transfer to unseen actors, natural conversations, or real-world dyadic interaction. This does not undermine the value of the dataset or the baseline as a proof-of-concept, but it does constrain the generalization claims in Sec. 8.
minor comments (6)
- [Sec. 3.1.1] The capture description says every scenario uses 'a pair of male and female actors,' but Sec. 7.1's actor-demographics limitation does not mention this restriction. Since same-gender and non-binary dyads are absent, the paper should explicitly acknowledge this in the limitations.
- [Sec. 3.2.5] The paper states that 'less than 5% of data' have hand anomalies (label swaps, finger extension errors, flickering) but does not list which sequences are affected. For a released dataset, per-sequence quality flags or a list of affected files would be important for users.
- [Sec. 4.1] Typo: 'phoenemes' should be 'phonemes.' Also, the capitalization of 'Diffspeaker' is inconsistent with 'DiffSpeaker' in the same section.
- [Table 1] The column header '#Ppl.' is ambiguous. For Ng et al. it reads 4, which the text elsewhere describes as two actors per session; please clarify whether this is total participants or participants per sequence.
- [Sec. 6] The statement 'all data is anonymized' is too strong: audio and motion capture of identifiable individuals are difficult to fully anonymize. Clarify the de-identification procedure (e.g., whether voice is modified, whether identity is removed from filenames/metadata).
- [Sec. 2.2] There is a stray comma and extra period in 'Most recently,. Xu et al. [2024]'.
Circularity Check
No circular derivation found: InterAct is a captured measurement, and the baseline is trained and tested on held-out splits; the acted-data representativeness concern is a stated limitation, not a circular step.
full rationale
The paper's central contribution is a newly captured multimodal dataset, not a derived quantity. The claim that InterAct contains diverse dynamic two-person interactions is supported by descriptive statistics (entropy, variance, contact rates, position/orientation scatter) computed directly from the captured motion, and by qualitative examples; these are measurements, not predictions reduced from inputs. The baseline method is trained on a training split (208 sequences) and evaluated on a held-out test split (33 sequences), a standard protocol; no fitted parameter is renamed as a prediction. The additional lip-accurate face dataset is a separate capture (one actor reading TIMIT sentences) used only for fine-tuning and is not the test set, so the reported lip-accuracy improvement is not forced by construction. The paper's own Sec. 7.1 concedes that the acted nature of the data 'may be regarded as less realistic' — this is an external-validity limitation, not a circularity. Prior architectures (DiffSpeaker, FaceFormer, LDA) are cited as building blocks rather than as authority for the dataset's novelty, and no uniqueness or forced-choice argument rests on a self-citation. No equation in the paper defines its target in terms of itself, and no evaluation metric is equivalent to a training objective by construction. The only substantive risk is whether acted performances generalize to real dyadic behavior, which the paper explicitly flags and which is a matter of empirical validity rather than circular derivation.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Temporal synchronization between VICON body capture and iPhone face/audio systems is frame-accurate via the timecode broadcast
- domain assumption Acted performances by 7 actors are representative of realistic daily dyadic interactions
- domain assumption VisualVoice speech separation extracts clean per-actor audio from the two-person recordings
- domain assumption MoCap marker and ARKit blendshape fits are accurate enough for downstream use
- ad hoc to paper The per-frame action labels C among {sit, walk, stand} computed heuristically are correct
Cite this review
Pith. "Pith review of InterAct: A Large-Scale Dataset of Dynamic, Expressive and Interactive Activities between Two People in Daily Scenarios." pith.science (2026). https://pith.science/paper/5B4FDJ6B
@misc{pith2026250905747,
author = {Pith},
title = {Pith review of: InterAct: A Large-Scale Dataset of Dynamic, Expressive and Interactive Activities between Two People in Daily Scenarios},
year = {2026},
howpublished = {\url{https://pith.science/paper/5B4FDJ6B}},
note = {Machine review of arXiv:2509.05747}
}
read the original abstract
We address the problem of accurate capture of interactive behaviors between two people in daily scenarios. Most previous works either only consider one person or solely focus on conversational gestures of two people, assuming the body orientation and/or position of each actor are constant or barely change over each interaction. In contrast, we propose to simultaneously model two people's activities, and target objective-driven, dynamic, and semantically consistent interactions which often span longer duration and cover bigger space. To this end, we capture a new multi-modal dataset dubbed InterAct, which is composed of 241 motion sequences where two people perform a realistic and coherent scenario for one minute or longer over a complete interaction. For each sequence, two actors are assigned different roles and emotion labels, and collaborate to finish one task or conduct a common interaction activity. The audios, body motions, and facial expressions of both persons are captured. InterAct contains diverse and complex motions of individuals and interesting and relatively long-term interaction patterns barely seen before. We also demonstrate a simple yet effective diffusion-based method that estimates interactive face expressions and body motions of two people from speech inputs. Our method regresses the body motions in a hierarchical manner, and we also propose a novel fine-tuning mechanism to improve the lip accuracy of facial expressions. To facilitate further research, the data and code is made available at https://hku-cg.github.io/interact/ .
Figures
Reference graph
Works this paper leans on
-
[3]
From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations. InArXiv. Evonne Ng, Sanjay Subramanian, Dan Klein, Angjoo Kanazawa, Trevor Darrell, and Shiry Ginosar. 2023. Can Language Models Learn to Listen?. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 10083–10093. Kunkun Pang, Dafei Qin, Yingruo Fan, Juli...
Pith/arXiv arXiv 2023
-
[2023]
InProceedings of the 25th international conference on multimodal interaction
The GENEA Challenge 2023: A large-scale evaluation of gesture generation models in monadic and dyadic settings. InProceedings of the 25th international conference on multimodal interaction. 792–801. Gilwoo Lee, Zhiwei Deng, Shugao Ma, Takaaki Shiratori, Siddhartha S Srinivasa, and Yaser Sheikh. 2019. Talking with hands 16.2 m: A large-scale dataset of syn...
Pith/arXiv arXiv 2023
-
[2024]
InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Convofusion: Multi-modal conversational diffusion for co-speech gesture synthesis. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1388–1398. InterAct: Dataset of Interactive Activities between Two People in Daily Scenarios 19 Tomohiko Mukai and Shigeru Kuriyama. 2005. Geostatistical motion interpolation. InACM SIGGRAP...
work page 2005
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.