Pith. sign in

REVIEW 4 major objections 6 minor 3 references

This paper introduces InterAct, a multimodal dataset of 241 minute-or-longer improvised scenes in which two actors perform realistic, emotionally labelled interactions while their audio, body motion, and facial expressions are captured simu

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-05 05:04 UTC pith:5B4FDJ6B

load-bearing objection InterAct is a genuinely useful dataset—first long-form dyadic dynamic interaction corpus with body, face, audio, and labels—but the paper overstates its uniqueness and the acting-realism gap is real, though honestly conceded. the 4 major comments →

arxiv 2509.05747 v1 pith:5B4FDJ6B submitted 2025-09-06 cs.CV cs.AIcs.LGcs.MAcs.RO

InterAct: A Large-Scale Dataset of Dynamic, Expressive and Interactive Activities between Two People in Daily Scenarios

classification cs.CV cs.AIcs.LGcs.MAcs.RO
keywords two-person interaction datasetmotion capturespeech-driven motion synthesisfacial animationdiffusion modeldyadic interactionmultimodal datasetinteractive behavior modeling
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper sets out to fix a gap in speech-to-motion research: existing datasets either capture one person, or capture two people only in static standing conversations. InterAct records 241 complete interactions between two actors, each lasting a minute or more, with synchronized audio, full-body marker motion, and facial meshes, under 26 emotion labels and 25 relationship categories. The authors argue this is the first dataset to combine long duration, dynamic movement across a large capture space, and simultaneous facial, body, and audio capture for two people. They also provide a speech-driven baseline that generates both actors' facial expressions and hierarchical body motions, showing the data can support interactive motion synthesis.

Core claim

The central claim is that daily two-person interaction—not just conversation—can be captured as a single synchronized multimodal record, and that this changes what speech-driven generation can attempt. InterAct is built by giving pairs of actors a role setup and an emotion label and letting them improvise a complete task or scene for a minute or longer in a 5m x 5m capture volume; 241 such sequences were collected from 7 actors. Each sequence ships separated speech for both actors, 53 body markers plus 20 finger markers per actor, facial mesh sequences converted to ARKit blendshapes, frame-level action labels, and relationship and emotion metadata. Statistical analysis shows wider spreads of

What carries the argument

The load-bearing object is the InterAct capture pipeline and its annotation scheme: 28 VICON cameras tracking 53 body and 20 finger markers per actor, head-mounted iPhones recording face RGB-D and speech, wireless timecode synchronization between the motion-capture system and the phones, and a speech-separation step that yields clean per-actor audio. The scenario prompts, which pair a relationship with an emotion, are what make the data objective-driven and semantically coherent rather than free-form chatting. The baseline's two mechanisms—hierarchical diffusion that first regresses lower-body joints and then upper-body joints conditioned on them, and denoiser-weight swapping during face fin

Load-bearing premise

The dataset's value assumes that one-minute improvised performances by seven actors, prompted with roles and emotions, are representative enough of real daily two-person interactions that models trained on them carry over to natural behavior.

What would settle it

Train the released baseline on InterAct and evaluate it on unscripted natural dyads captured with ordinary video and pose estimation; if the generated relative-distance and orientation dynamics match those of static-conversation baselines rather than the natural dyads, or if human raters cannot distinguish InterAct-trained outputs from outputs trained on static conversational data, the acting-data representativeness claim is falsified.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Audio-driven generation can now target two people who walk, sit, shift position, and orient toward or away from each other, not just speakers rooted in place.
  • The dataset supplies a benchmark in which models must reproduce long-range interaction patterns, proxemics, and relative body orientation, not only local gestures.
  • The hierarchical body diffusion recipe suggests a practical way to split high-dimensional two-person motion generation into lower-body control and upper-body detail.
  • The denoiser-swap fine-tuning trick gives a cheap way to improve lip sync on a small high-quality face dataset without retraining the whole model.
  • Because each sequence carries emotion, relationship, and action labels, generated motion can be steered by those labels at inference time.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If improvisation under role prompts generalizes, the same capture protocol could be used with non-professional participants or scripted naturalistic tasks; a useful check would be comparing InterAct's relative-distance and orientation distributions against those of unscripted dyadic video.
  • The lip-accuracy denoiser swap is separable from the rest of the face model in spirit; testing it inside a different speech-to-face architecture would show whether the trick transfers.
  • The captured hand-contact statistics, such as higher hand-hand self-contact rates in friends and social scenarios, suggest InterAct could serve as a source for proxemics and nonverbal-communication studies, though the paper only reports the statistics and does not develop that analysis.
  • The flat-ground, prop-free setup means models trained on the dataset may not capture object-mediated interactions; augmenting with object interaction data or synthetic props would be a natural test of transfer.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces InterAct, a two-person multimodal motion-capture dataset: 241 sequences (about 10 hours) of two actors improvising role- and emotion-prompted daily-life scenarios, with synchronized audio, body MoCap (53 markers plus finger markers), and facial ARKit blendshapes for both participants. The authors also propose a speech-conditioned diffusion baseline that hierarchically generates two people's body motions and facial expressions, and they evaluate it on held-out InterAct sequences with quantitative metrics and user studies. The dataset and code are promised for public release.

Significance. If the claims are corrected and the data are released, InterAct would be a useful resource for dyadic interaction research: it is one of the few datasets with simultaneous body, face, and audio capture of two interacting people in non-stationary scenarios, and it includes long (1+ minute) coherent activities. The capture pipeline is described in detail, statistical comparisons are provided, and the baseline is a reproducible attempt at a genuinely hard joint generation task. I agree with the paper that evaluating the baseline on held-out sequences from the same dataset is standard practice and not circular; however, the paper's novelty claims are currently overstated, and several external-validity claims go beyond what the evidence supports.

major comments (4)
  1. [Sec. 3.2 and Table 1] The claim that InterAct is 'the only one that captures both actors' body motions and facial expressions at the same time' is directly contradicted by the paper's own Table 1: the Ng et al. (2024) row lists Audio ✓, Hands Detailed, Face ✓, and Coherence ✓. Since the related-work section also describes Ng et al. as a two-person conversation dataset, the uniqueness claim must be narrowed to the specific combination of dynamic, large-range, two-person interaction with simultaneous face capture, or the comparison table must be corrected. This is load-bearing because the abstract and introduction frame InterAct's contribution as a first-of-kind resource.
  2. [Sec. 3.2 and Fig. 4 caption] The main text says Ng et al. (2024) 'records both actors primarily during conversation standing still,' but the Fig. 4 caption states that 'only one actor is captured in BEAT and in [Ng et al. 2024], so analysis in the lower part is not applicable on them.' These statements are mutually inconsistent and affect the central statistical comparison: whether InterAct uniquely enables two-person relative-position/orientation analysis depends on whether Ng et al. actually captured both participants' body motion. The authors must resolve this contradiction and, if both actors are available in Ng et al., include the lower-panel comparison or justify its exclusion.
  3. [Sec. 7.1 and Sec. 8] The dataset's value for 'human behavior grounding' and 'real-life, physically-informed data' (Sec. 8) rests on the assumption that improvised performances by 7 actors on a flat 5m x 5m stage, without props and sometimes with adults playing children, are representative of natural dyadic behavior. Section 7.1 concedes the acting nature but no validation is offered: the statistical analyses in Sec. 3.2 measure internal diversity, not fidelity to real interaction. I do not consider the use of acted data fatal by itself, since similar resources (BEAT, Ng et al.) are also acted or controlled, but the conclusion's wording overstates ecological validity. The authors should either provide an external comparison (e.g., proxemics or contact statistics against a naturalistic corpus such as TalkHands) or substantially soften the 'real-life' and 'socially-aware robotics' claims.
  4. [Sec. 5] All baseline evaluations, including Tables 2 and 4 and both user studies, are performed on a held-out portion of InterAct itself, i.e., in-distribution with the same actors and the same improvised acting style. The paper should state this limitation explicitly when claiming the method is 'effective' and avoid implying that generated motions transfer to unseen actors, natural conversations, or real-world dyadic interaction. This does not undermine the value of the dataset or the baseline as a proof-of-concept, but it does constrain the generalization claims in Sec. 8.
minor comments (6)
  1. [Sec. 3.1.1] The capture description says every scenario uses 'a pair of male and female actors,' but Sec. 7.1's actor-demographics limitation does not mention this restriction. Since same-gender and non-binary dyads are absent, the paper should explicitly acknowledge this in the limitations.
  2. [Sec. 3.2.5] The paper states that 'less than 5% of data' have hand anomalies (label swaps, finger extension errors, flickering) but does not list which sequences are affected. For a released dataset, per-sequence quality flags or a list of affected files would be important for users.
  3. [Sec. 4.1] Typo: 'phoenemes' should be 'phonemes.' Also, the capitalization of 'Diffspeaker' is inconsistent with 'DiffSpeaker' in the same section.
  4. [Table 1] The column header '#Ppl.' is ambiguous. For Ng et al. it reads 4, which the text elsewhere describes as two actors per session; please clarify whether this is total participants or participants per sequence.
  5. [Sec. 6] The statement 'all data is anonymized' is too strong: audio and motion capture of identifiable individuals are difficult to fully anonymize. Clarify the de-identification procedure (e.g., whether voice is modified, whether identity is removed from filenames/metadata).
  6. [Sec. 2.2] There is a stray comma and extra period in 'Most recently,. Xu et al. [2024]'.

Circularity Check

0 steps flagged

No circular derivation found: InterAct is a captured measurement, and the baseline is trained and tested on held-out splits; the acted-data representativeness concern is a stated limitation, not a circular step.

full rationale

The paper's central contribution is a newly captured multimodal dataset, not a derived quantity. The claim that InterAct contains diverse dynamic two-person interactions is supported by descriptive statistics (entropy, variance, contact rates, position/orientation scatter) computed directly from the captured motion, and by qualitative examples; these are measurements, not predictions reduced from inputs. The baseline method is trained on a training split (208 sequences) and evaluated on a held-out test split (33 sequences), a standard protocol; no fitted parameter is renamed as a prediction. The additional lip-accurate face dataset is a separate capture (one actor reading TIMIT sentences) used only for fine-tuning and is not the test set, so the reported lip-accuracy improvement is not forced by construction. The paper's own Sec. 7.1 concedes that the acted nature of the data 'may be regarded as less realistic' — this is an external-validity limitation, not a circularity. Prior architectures (DiffSpeaker, FaceFormer, LDA) are cited as building blocks rather than as authority for the dataset's novelty, and no uniqueness or forced-choice argument rests on a self-citation. No equation in the paper defines its target in terms of itself, and no evaluation metric is equivalent to a training objective by construction. The only substantive risk is whether acted performances generalize to real dyadic behavior, which the paper explicitly flags and which is a matter of empirical validity rather than circular derivation.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

The central claim is a captured dataset, so the ledger contains no fitted free parameters and no invented entities. The key burdens are representativeness of acted data, capture fidelity, speech separation quality, and the correctness of the heuristic action labels used by the baseline.

axioms (5)
  • domain assumption Temporal synchronization between VICON body capture and iPhone face/audio systems is frame-accurate via the timecode broadcast
    Sec. 3.1.1 states a script synchronizes recording 'at frame-level accuracy'; multi-modal alignment is assumed, with no reported validation against an independent clock.
  • domain assumption Acted performances by 7 actors are representative of realistic daily dyadic interactions
    Sec. 3.1.1 (Diversity) and Sec. 7.1 (Nature and Scope of Performances): actors improvise from role/emotion prompts; the paper admits the acting nature and imperfect demographics. This assumption underlies the dataset's external validity.
  • domain assumption VisualVoice speech separation extracts clean per-actor audio from the two-person recordings
    Sec. 3.1.1 applies VisualVoice to separate overlapping speech; separation accuracy is not evaluated on this capture setup, so audio-condition quality is unverified.
  • domain assumption MoCap marker and ARKit blendshape fits are accurate enough for downstream use
    Sec. 3.1.1 and Sec. 3.2.5: body markers are exported to BVH and face meshes to ARKit via least-squares; the paper notes less than 5% hand anomalies and small face conversion errors, but does not quantify errors for the remaining data.
  • ad hoc to paper The per-frame action labels C among {sit, walk, stand} computed heuristically are correct
    Sec. 3.1.1 and Sec. 4.2 introduce the label as computed 'in a heuristic way'; generation quality depends on it, yet no accuracy measure is reported.

pith-pipeline@v1.3.0-alltime-deepseek · 21444 in / 12631 out tokens · 134920 ms · 2026-08-05T05:04:26.306729+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of InterAct: A Large-Scale Dataset of Dynamic, Expressive and Interactive Activities between Two People in Daily Scenarios." pith.science (2026). https://pith.science/paper/5B4FDJ6B

@misc{pith2026250905747,
  author       = {Pith},
  title        = {Pith review of: InterAct: A Large-Scale Dataset of Dynamic, Expressive and Interactive Activities between Two People in Daily Scenarios},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5B4FDJ6B}},
  note         = {Machine review of arXiv:2509.05747}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We address the problem of accurate capture of interactive behaviors between two people in daily scenarios. Most previous works either only consider one person or solely focus on conversational gestures of two people, assuming the body orientation and/or position of each actor are constant or barely change over each interaction. In contrast, we propose to simultaneously model two people's activities, and target objective-driven, dynamic, and semantically consistent interactions which often span longer duration and cover bigger space. To this end, we capture a new multi-modal dataset dubbed InterAct, which is composed of 241 motion sequences where two people perform a realistic and coherent scenario for one minute or longer over a complete interaction. For each sequence, two actors are assigned different roles and emotion labels, and collaborate to finish one task or conduct a common interaction activity. The audios, body motions, and facial expressions of both persons are captured. InterAct contains diverse and complex motions of individuals and interesting and relatively long-term interaction patterns barely seen before. We also demonstrate a simple yet effective diffusion-based method that estimates interactive face expressions and body motions of two people from speech inputs. Our method regresses the body motions in a hierarchical manner, and we also propose a novel fine-tuning mechanism to improve the lip accuracy of facial expressions. To facilitate further research, the data and code is made available at https://hku-cg.github.io/interact/ .

Figures

Figures reproduced from arXiv: 2509.05747 by Dafei Qin, Junichi Yamagishi, Leo Ho, Mingyi Shi, Taku Komura, Wangpok Tse, Wei Liu, Yinghao Huang.

Figure 1
Figure 1. Figure 1: We capture and model realistic interactions between two people in daily life. Our dataset includes [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Breakdown of InterAct dataset by meta-relationships, meta-emotions, and gender. Fine-grained [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Two pairs of actors during capture sessions. Their audios, facial expressions and body motions are all [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Statistical comparisons between InterAct and previous datasets w.r.t. individual motions (upper part) [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Left: We calculate entropy values for body motions of different actors, meta-relationships, and meta [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Pipeline of our system. Given raw audio signals of two people, BERT feature, relative orientation [PITH_FULL_IMAGE:figures/full_fig_p011_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Result of face user study. Users are asked to [PITH_FULL_IMAGE:figures/full_fig_p013_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Top: Sample of our results where two actors are having a conversation. Bottom: Skeleton comparison [PITH_FULL_IMAGE:figures/full_fig_p014_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Left: Base model, Middle: Final model with denoiser weight injection, Right: Fine-tuned model. All [PITH_FULL_IMAGE:figures/full_fig_p026_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Interface of qualitative study for body motions. [PITH_FULL_IMAGE:figures/full_fig_p026_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

3 extracted references · 1 canonical work pages

  1. [3]

    From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations. InArXiv. Evonne Ng, Sanjay Subramanian, Dan Klein, Angjoo Kanazawa, Trevor Darrell, and Shiry Ginosar. 2023. Can Language Models Learn to Listen?. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 10083–10093. Kunkun Pang, Dafei Qin, Yingruo Fan, Juli...

  2. [2023]

    InProceedings of the 25th international conference on multimodal interaction

    The GENEA Challenge 2023: A large-scale evaluation of gesture generation models in monadic and dyadic settings. InProceedings of the 25th international conference on multimodal interaction. 792–801. Gilwoo Lee, Zhiwei Deng, Shugao Ma, Takaaki Shiratori, Siddhartha S Srinivasa, and Yaser Sheikh. 2019. Talking with hands 16.2 m: A large-scale dataset of syn...

  3. [2024]

    InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Convofusion: Multi-modal conversational diffusion for co-speech gesture synthesis. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1388–1398. InterAct: Dataset of Interactive Activities between Two People in Daily Scenarios 19 Tomohiko Mukai and Shigeru Kuriyama. 2005. Geostatistical motion interpolation. InACM SIGGRAP...