REVIEW 4 major objections 6 minor 54 references
MIME, a dedicated encoder for two-person interactive motion, aligns text with motion by modeling each actor separately and their interaction explicitly; the paper reports consistent retrieval gains over fusion baselines and improved semanti
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 19:27 UTC pith:J33QZBZK
load-bearing objection Solid first dedicated dyadic text-motion retrieval encoder with honest ablations; the retrieval result holds, but the downstream transfer claim is not isolated because only the text encoder is used. the 4 major comments →
MIME: Multimodal Interactive Motion Encoder
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Modeling a two-person interaction as two coupled streams, not one concatenated sequence, yields a more discriminative shared text-motion space. MIME keeps actor streams separate, injects explicit relational features, and exchanges information via bidirectional co-attention. With contrastive training and a hardness-increasing curriculum, it reports text-to-motion R@1 of 38.87 at 500 samples and 20.67 at 2,000, versus 34.64 and 18.32 for the strongest early-fusion baseline. Frozen as a prior, it raises R@1 on unseen InterHuman from 0.4732 to 0.4924 (TIMotion) and 0.4372 to 0.4489 (InterMask); FID holds.
What carries the argument
The load-bearing object is the stream-based bidirectional co-attention transformer with explicit relational feature injection. Each actor's frame feature is augmented with root-pair distance, signed root displacement, feature difference, and a scalar rho for relative motion magnitude, then projected and added to both streams. Layers then apply person-specific self-attention, bidirectional cross-attention over full sequences, and feed-forward updates, followed by frame-wise fusion and attention pooling. A curriculum sampler shifts negatives from distant to close over ten epochs, sharpening sensitivity to fine-grained interaction differences. The architecture preserves actor identity while mak
Load-bearing premise
The comparison assumes that the strong pretrained baseline, LaMP, is fairly tested through only its motion backbone under a shared early-fusion protocol; if LaMP's full pretraining is essential to its representation, the reported margin over LaMP overstates MIME's advantage.
What would settle it
Run the retrieval benchmark with LaMP's complete pretrained pipeline (its own pretraining and native evaluation) instead of only its backbone under early fusion; if it reaches or exceeds MIME's 20.67 text-to-motion R@1 at the 2,000-sample gallery, the claimed architectural gain is a protocol artifact.
If this is right
- If MIME's representation is as discriminative as reported, text-to-motion retrieval for two-person interactions can be served by a single frozen encoder rather than per-task fusion modules.
- The frozen MIME prior improves semantic alignment metrics in TIMotion without degrading FID, suggesting interaction-aware conditioning can be added to existing generators at low cost.
- The gains grow with gallery size and with semantically similar distractors, implying the benefit is strongest precisely where retrieval is hardest.
- Ablations show explicit relational features carry the most consistent contribution, so future interaction encoders should treat relative geometry as first-class input, not something to be learned implicitly.
- Compact variants of MIME with fewer motion-side parameters than a late-fusion baseline still beat 16 of 18 baseline results, indicating the architecture rather than parameter count drives the improvement.
Where Pith is reading between the lines
- The paper does not claim, but the same embedding could serve as a mining tool for building editing pairs: nearby motions that preserve an interaction while differing in one describable attribute, which the qualitative examples gesture toward.
- The co-attention-plus-relational-features recipe is likely portable to larger groups by adding streams and pairwise relational features, though only dyads are evaluated.
- A testable extension the paper leaves implicit: if the curriculum is what buys fine-grained discrimination, the same hardness schedule should improve other multimodal alignment tasks, such as audio-motion or text-scene retrieval, especially in their hard tails.
- The InterMask FID increase suggests the way a frozen prior is fused into a generator matters; a lightweight adapter can trade distribution fidelity for semantic alignment, so integration design is a variable, not a given.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. MIME is a text-motion representation-learning method for two-person interactions. The architecture keeps separate actor streams, injects explicit frame-level relational features, applies bidirectional co-attention between the streams, and trains with a symmetric contrastive objective plus a curriculum batch sampler. Retrieval experiments on Inter-X compare MIME against TMR early/late fusion and a LaMP backbone, reporting consistent gains that grow with gallery size (e.g., text-to-motion R@1 20.67 vs. 18.32 for TMR Early Fusion at a 2,000-sample gallery). The paper also evaluates MIME as a frozen auxiliary text-conditioning prior in TIMotion and InterMask on the unseen InterHuman dataset, reporting improved retrieval-based alignment metrics with comparable FID in TIMotion. The central claim is that an interaction-aware multimodal encoder yields a more discriminative shared text-motion representation for dyadic interaction retrieval and transfer.
Significance. If the central search claim holds, MIME is a useful reusable representation for two-person interaction retrieval and conditioning, filling a gap left by single-person encoders. The paper has several genuine strengths: the retrieval evaluation uses a held-out test gallery and a second unseen dataset with the standard InterHuman evaluator; compact-variant experiments (Tables 4-5) directly address the capacity confound; hard-gallery experiments (Table 6) and full rank-distribution analysis (Fig. 6) provide evidence beyond a single R@1 number; and the ablations isolate components. The main load-bearing weakness is that the downstream transfer experiment does not actually exercise the proposed motion encoder, so the transfer claim is not isolated from the effect of fine-tuning the CLIP text encoder on interaction data. This, together with the questionable LaMP baseline protocol and the lack of significance testing in Table 3, means the paper currently overstates the breadth of its validation.
major comments (4)
- [Sec. 4.8, Table 3] The downstream validation does not test the proposed interaction-aware motion encoder. The text says MIME is integrated by projecting "MIME's interaction-aware text representation" through a linear-GELU-linear adapter, i.e., only the text encoder output is used; the co-attention motion encoder, explicit relational features, and person-type embeddings are not exercised at inference. Since the text encoder is a CLIP encoder fine-tuned on Inter-X with a contrastive objective, the observed R@1 gains (TIMotion 0.4732→0.4924; InterMask 0.4372→0.4489) could be entirely due to text-side adaptation. A control prior using the text encoder of an early-fusion baseline trained under the same protocol is needed, or the claim should be restricted to "fine-tuned interaction-specific text representation" rather than the MIME architecture. As written, the title claim that MIME's representation transfers i
- [Sec. 4.4, Table 1] The LaMP comparison is not apples-to-apples. The paper states that LaMP is not run with its own pretraining or generation pipeline but instead uses only its motion backbone under the same early-fusion retrieval protocol. This may understate LaMP if its native pretraining is important for representation quality. Because the abstract's headline is against early/late fusion baselines, the central retrieval claim does not collapse, but the text in Sec. 4.5 and the user study treat LaMP as a comparable baseline. Either run LaMP with its native pretrained representation, or explicitly relabel this row as "LaMP backbone (re-trained)" and remove the unqualified "outperforms LaMP" language.
- [Table 3] The downstream improvements are reported as mean ± 95% CI, but no significance tests are given. Several differences are small relative to the reported intervals: TIMotion R@1 0.4732±0.0141 vs. 0.4924±0.0107, R@2 0.6312±0.0120 vs. 0.6404±0.0119; InterMask R@1 0.4372±0.0117 vs. 0.4489±0.0117. These intervals overlap, so the claim that MIME "improves semantic alignment" should be backed by a paired significance test (e.g., bootstrap or permutation) or stated as a trend. This is load-bearing because the cross-dataset transfer claim is one of the paper's two main contributions.
- [Sec. 4.6, Supp. C] The user study reports large relative improvements (42% more aligned, etc.) from 30 participants on 10 queries, but no error bars, inter-rater variability, or significance testing are provided. This is qualitative support rather than evidence, and the summary in the main text overstates it. The supplementary should report per-criterion means with confidence intervals and a paired test.
minor comments (6)
- [Sec. 4.7, Table 2] At 500- and 1,000-sample galleries, the no-curriculum and no-co-attention variants occasionally match or exceed the full model (e.g., M2T R@1 at 500: –Cur 48.70 vs. Full 47.10). The text handles this by attributing the benefit to harder retrieval, but it would help to state explicitly that the component gains are not monotonic across gallery sizes and to discuss the variance.
- [Sec. 3.2.2] Minor notation: in Eqs. (3)-(4), the augmented feature includes a scalar distance and a 3D displacement, so the dimension 139 is correct, but the typesetting of the norm and vector is easy to misread. Consider adding explicit dimensions to the equations.
- [Supp. A] The supplementary says code and checkpoints "will be made open source and publicly available upon paper acceptance." For a representation-learning paper whose contribution is a reusable encoder, providing these at submission time (or in a reproducible appendix) would strengthen the claims.
- [Sec. 4.3] The definition of MM Dist is given only as "average distance between matched text and motion embeddings." Since the downstream table reports MM Dist for generated motions, specify which embedding space and which normalizations are used (e.g., the original generators' metric spaces).
- [Fig. 1 caption] Typo: "both streamsh a" should read "both streams h_a".
- [Sec. 4.5] The sentence "For motion to text retrieval, MIME improves over the strongest baseline at R@1, R@3, and R@5 by 4.3%, 7.4%, and 5.3% respectively" should specify the gallery size and which baseline it is computed against, since values differ across Table 1 rows.
Circularity Check
No significant circularity: retrieval and downstream claims rest on held-out evaluations and external benchmarks, not on the paper's own fitted parameters or self-citations.
full rationale
The paper's central retrieval claim is an empirical result on held-out Inter-X test galleries (Secs. 4.2–4.5): MIME's R@1 improvements are measured against fixed baseline models using the same deterministic subsets and a standard contrastive objective, not derived from MIME's own equations or from a fitted parameter renamed as a prediction. The downstream validation (Sec. 4.8) uses the unseen InterHuman dataset and the standard InterHuman evaluator, so the reported MM Distance, R-precision, and FID values are externally computed quantities rather than consequences of MIME's training objective. The architectural components—relational features, co-attention, curriculum sampling—are input features and training regularizers, not definitions of the evaluation metrics. The self-citations ([9], [10], [17]) appear only as related-work context and are not load-bearing; the TMR objective [29] and curriculum sampler [42] are external published methods. The LaMP baseline is configured under an early-fusion protocol, which may be an unfair or incomplete comparison, but that is a baseline-construction concern, not circularity. Similarly, the downstream experiment uses only MIME's fine-tuned text representation through a lightweight adapter, so it cannot isolate the motion-side co-attention contribution; this is an attribution/control limitation, not a circular derivation, because the reported gains are not equivalent to the paper's inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (3)
- Curriculum max hardness eta_max =
0.25
- Curriculum warm-up and ramp epochs =
3 warm-up + 10 ramp
- Architecture hyperparameters =
D=512, L_c=4, heads=4, dropout=0.1, batch 128, lr 1e-4
axioms (5)
- domain assumption CLIP text representations are a suitable starting point for interaction captions and can be fine-tuned to motion semantics.
- domain assumption The 135-D SMPL-derived per-frame pose/root features contain enough information to infer interaction semantics (contact, synchronization, relative geometry).
- domain assumption Inter-X and InterHuman captions are reliable ground truth for retrieval and evaluation.
- ad hoc to paper Linear differences of 6D rotation vectors are useful explicit relational inputs.
- ad hoc to paper The curriculum batch sampler from Wu et al. transfers from deep embedding learning to text-motion contrastive training.
invented entities (3)
-
Person-type embeddings e_a, e_b
no independent evidence
-
Relation-type embedding e_r
no independent evidence
-
Learnable temporal pooling query q
no independent evidence
read the original abstract
Text-motion representation learning has advanced rapidly, with growing interest in multi person interactions for animation, AR/VR, and embodied AI. These settings require representations that align language with both individual actor dynamics and the relationships between actors. We introduce the Multimodal Interactive Motion Encoder (MIME), which, to our knowledge, represents the first dedicated multimodal encoder designed specifically for two person interactive motion. MIME captures individual and shared structure using stream based co-attention with explicit interaction features and curriculum based contrastive training. On Inter-X text-motion retrieval, MIME consistently outperforms early and late fusion baselines across gallery sizes, achieving a 12.8% relative improvement in text-to-motion R@1 at a 2,000-sample gallery. We further evaluate MIME as a frozen auxiliary prior within TIMotion and InterMask on the unseen InterHuman dataset. MIME improves semantic alignment metrics while maintaining comparable FID in TIMotion. These results show that interaction aware multimodal encoding improves multi person motion retrieval and transfers across datasets to support downstream motion generation.
Figures
Reference graph
Works this paper leans on
-
[1]
Motionfix: Text-driven 3d human motion editing
Nikos Athanasiou, Alp ´ar Cseke, Markos Diomataris, Michael J Black, and G ¨ul Varol. Motionfix: Text-driven 3d human motion editing. InSIGGRAPH Asia 2024 Conference Papers, pages 1–11, 2024. 2, 11
2024
-
[2]
Make-an-animation: Large-scale text- conditional 3d human motion generation
Samaneh Azadi, Akbar Shah, Thomas Hayes, Devi Parikh, and Sonal Gupta. Make-an-animation: Large-scale text- conditional 3d human motion generation. InProceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 15039–15048, 2023. 1
2023
-
[3]
Curriculum learning
Yoshua Bengio, J ´erˆome Louradour, Ronan Collobert, and Ja- son Weston. Curriculum learning. InProceedings of the 26th 8 annual international conference on machine learning, pages 41–48, 2009. 1
2009
-
[4]
3d human interaction generation: A survey.arXiv preprint arXiv:2503.13120, 2025
Siyuan Fan, Wenke Huang, Xiantao Cai, and Bo Du. 3d human interaction generation: A survey.arXiv preprint arXiv:2503.13120, 2025. 1
Pith/arXiv arXiv 2025
-
[5]
Remos: 3d motion- conditioned reaction synthesis for two-person interactions
Anindita Ghosh, Rishabh Dabral, Vladislav Golyanik, Chris- tian Theobalt, and Philipp Slusallek. Remos: 3d motion- conditioned reaction synthesis for two-person interactions. InEuropean conference on computer vision, pages 418–437. Springer, 2024. 2
2024
-
[6]
Ac- tion2motion: Conditioned generation of 3d human motions
Chuan Guo, Xinxin Zuo, Sen Wang, Shihao Zou, Qingyao Sun, Annan Deng, Minglun Gong, and Li Cheng. Ac- tion2motion: Conditioned generation of 3d human motions. InProceedings of the 28th ACM international conference on multimedia, pages 2021–2029, 2020. 5
2021
-
[7]
Generating diverse and natural 3d human motions from text
Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. Generating diverse and natural 3d human motions from text. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5152–5161, 2022. 2, 5
2022
-
[8]
Momask: Generative masked model- ing of 3d human motions
Chuan Guo, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang, and Li Cheng. Momask: Generative masked model- ing of 3d human motions. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1900–1910, 2024. 2
1900
-
[9]
Mdd: A dataset for text-and- music conditioned duet dance generation
Prerit Gupta, Jason Alexander Fotso-Puepi, Zhengyuan Li, Jay Mehta, and Aniket Bera. Mdd: A dataset for text-and- music conditioned duet dance generation. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 13932–13941, 2025. 2
2025
-
[10]
Prerit Gupta, Shourya Verma, Ananth Grama, and Aniket Bera. Unified multi-modal interactive & reactive 3d motion generation via rectified flow.arXiv preprint arXiv:2509.24099, 2025. 2
arXiv 2025
-
[11]
Gaussian error linear units (gelus).arXiv preprint arXiv:1606.08415, 2016
Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus).arXiv preprint arXiv:1606.08415, 2016. 4
Pith/arXiv arXiv 2016
-
[12]
Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017. 5
2017
-
[13]
Muhammad Gohar Javed, Chuan Guo, Li Cheng, and Xingyu Li. Intermask: 3d human interaction genera- tion via collaborative masked modeling.arXiv preprint arXiv:2410.10010, 2024. 8
Pith/arXiv arXiv 2024
-
[14]
Motiongpt: Human motion as a foreign lan- guage.Advances in Neural Information Processing Systems, 36:20067–20079, 2023
Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu, Gang Yu, and Tao Chen. Motiongpt: Human motion as a foreign lan- guage.Advances in Neural Information Processing Systems, 36:20067–20079, 2023. 2
2023
-
[15]
Two-in-one: Unified multi-person interactive motion gener- ation by latent diffusion transformer, 2024
Boyuan Li, Xihua Wang, Ruihua Song, and Wenbing Huang. Two-in-one: Unified multi-person interactive motion gener- ation by latent diffusion transformer, 2024. 1
2024
-
[16]
Duolando: Follower gpt with off-policy reinforcement learn- ing for dance accompaniment
Siyao Li, Tianpei Gu, Zhitao Yang, Zhengyu Lin, Ziwei Liu, Henghui Ding, Lei Yang, and Chen Change Loy. Duolando: Follower gpt with off-policy reinforcement learn- ing for dance accompaniment. InInternational Conference on Learning Representations, pages 810–829, 2024. 2
2024
-
[17]
Simmotionedit: Text-based human motion editing with motion similarity pre- diction
Zhengyuan Li, Kai Cheng, Anindita Ghosh, Uttaran Bhat- tacharya, Liangyan Gui, and Aniket Bera. Simmotionedit: Text-based human motion editing with motion similarity pre- diction. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 27827– 27837, 2025. 2
2025
-
[18]
Lamp: Language-motion pretraining for motion generation, retrieval, and captioning
Zhe Li, Weihao Yuan, Yisheng He, Lingteng Qiu, Shen- hao Zhu, Xiaodong Gu, Weichao Shen, Yuan Dong, Zi- long Dong, and Laurence Yang. Lamp: Language-motion pretraining for motion generation, retrieval, and captioning. InInternational Conference on Learning Representations, pages 84238–84250, 2025. 2
2025
-
[19]
Intergen: Diffusion-based multi-human motion genera- tion under complex interactions.International Journal of Computer Vision, 132(9):3463–3483, 2024
Han Liang, Wenqian Zhang, Wenxuan Li, Jingyi Yu, and Lan Xu. Intergen: Diffusion-based multi-human motion genera- tion under complex interactions.International Journal of Computer Vision, 132(9):3463–3483, 2024. 2, 5
2024
-
[20]
Matthew Loper, Naureen Mahmood, Javier Romero, Ger- ard Pons-Moll, and Michael J. Black. SMPL: A skinned multi-person linear model.ACM Trans. Graphics (Proc. SIGGRAPH Asia), 34(6):248:1–248:16, 2015. 3
2015
-
[21]
Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017. 5
Pith/arXiv arXiv 2017
-
[22]
Hierarchical question-image co-attention for visual question answering.Advances in neural information processing sys- tems, 29, 2016
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. Hierarchical question-image co-attention for visual question answering.Advances in neural information processing sys- tems, 29, 2016. 2
2016
-
[23]
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.Advances in neural information processing systems, 32, 2019
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.Advances in neural information processing systems, 32, 2019. 2
2019
-
[24]
Rethinking diffusion for text-driven human motion generation: Redundant representations, evaluation, and masked autoregression
Zichong Meng, Yiming Xie, Xiaogang Peng, Zeyu Han, and Huaizu Jiang. Rethinking diffusion for text-driven human motion generation: Redundant representations, evaluation, and masked autoregression. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 27859– 27871, 2025. 2
2025
-
[25]
Multimodal deep learn- ing
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, Andrew Y Ng, et al. Multimodal deep learn- ing. InIcml, pages 689–696, 2011. 2
2011
-
[26]
Uniegomo- tion: A unified model for egocentric motion reconstruction, forecasting, and generation
Chaitanya Patel, Hiroki Nakamura, Yuta Kyuragi, Kazuki Kozuka, Juan Carlos Niebles, and Ehsan Adeli. Uniegomo- tion: A unified model for egocentric motion reconstruction, forecasting, and generation. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 10318– 10329, 2025. 1
2025
-
[27]
Deepmimic: Example-guided deep reinforce- ment learning of physics-based character skills.ACM Trans- actions On Graphics (TOG), 37(4):1–14, 2018
Xue Bin Peng, Pieter Abbeel, Sergey Levine, and Michiel Van de Panne. Deepmimic: Example-guided deep reinforce- ment learning of physics-based character skills.ACM Trans- actions On Graphics (TOG), 37(4):1–14, 2018. 1
2018
-
[28]
Temos: Generating diverse human motions from textual descriptions
Mathis Petrovich, Michael J Black, and G ¨ul Varol. Temos: Generating diverse human motions from textual descriptions. InEuropean conference on computer vision, pages 480–497. Springer, 2022. 2
2022
-
[29]
Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis
Mathis Petrovich, Michael J Black, and G ¨ul Varol. Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 9488–9497, 2023. 2, 4, 11 9
2023
-
[30]
The kit motion-language dataset.Big data, 4(4):236–252,
Matthias Plappert, Christian Mandery, and Tamim Asfour. The kit motion-language dataset.Big data, 4(4):236–252,
-
[31]
Babel: Bodies, action and behavior with english la- bels
Abhinanda R Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez, and Michael J Black. Babel: Bodies, action and behavior with english la- bels. InProceedings of the IEEE/CVF conference on com- puter vision and pattern recognition, pages 722–731, 2021. 2
2021
-
[32]
Learning transferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. InInternational conference on machine learning, pages 8748–8763. PmLR, 2021. 3
2021
-
[33]
Kimodo: Scal- ing controllable human motion generation.arXiv preprint arXiv:2603.15546, 2026
Davis Rempe, Mathis Petrovich, Ye Yuan, Haotian Zhang, Xue Bin Peng, Yifeng Jiang, Tingwu Wang, Umar Iqbal, David Minor, Michael de Ruyter, et al. Kimodo: Scal- ing controllable human motion generation.arXiv preprint arXiv:2603.15546, 2026. 1
arXiv 2026
-
[34]
Junlong Ren, Gangjian Zhang, Honghao Fu, Pengcheng Wu, and Hao Wang. Wamo: Wavelet-enhanced multi- frequency trajectory analysis for fine-grained text-motion re- trieval.arXiv preprint arXiv:2508.03343, 2025. 2
Pith/arXiv arXiv 2025
-
[35]
A survey on human interaction mo- tion generation.International Journal of Computer Vision, 134(3):113, 2026
Kewei Sui, Anindita Ghosh, Inwoo Hwang, Bing Zhou, Jian Wang, and Chuan Guo. A survey on human interaction mo- tion generation.International Journal of Computer Vision, 134(3):113, 2026. 2
2026
-
[36]
Motionclip: Exposing human motion generation to clip space
Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or. Motionclip: Exposing human motion generation to clip space. InEuropean Conference on Com- puter Vision, pages 358–374. Springer, 2022. 2
2022
-
[37]
Human motion dif- fusion model.arXiv preprint arXiv:2209.14916, 2022
Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H Bermano. Human motion dif- fusion model.arXiv preprint arXiv:2209.14916, 2022. 2
Pith/arXiv arXiv 2022
-
[38]
Multimodal transformer for unaligned multimodal language sequences
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. Multimodal transformer for unaligned multimodal language sequences. InProceedings of the 57th annual meeting of the association for computational linguistics, pages 6558–6569,
-
[39]
Neural discrete representation learning.Advances in neural information pro- cessing systems, 30, 2017
Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning.Advances in neural information pro- cessing systems, 30, 2017. 2
2017
-
[40]
Cross-modal retrieval: a system- atic review of methods and future directions.Proceedings of the IEEE, 112(11):1716–1754, 2025
Tianshi Wang, Fengling Li, Lei Zhu, Jingjing Li, Zheng Zhang, and Heng Tao Shen. Cross-modal retrieval: a system- atic review of methods and future directions.Proceedings of the IEEE, 112(11):1716–1754, 2025. 1
2025
-
[41]
Timotion: Temporal and in- teractive framework for efficient human-human motion gen- eration
Yabiao Wang, Shuo Wang, Jiangning Zhang, Ke Fan, Jiafu Wu, Zhucun Xue, and Yong Liu. Timotion: Temporal and in- teractive framework for efficient human-human motion gen- eration. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025. 1, 2
2025
-
[42]
Sampling matters in deep embed- ding learning
Chao-Yuan Wu, Raghavan Manmatha, Alexander J Smola, and Philipp Krahenbuhl. Sampling matters in deep embed- ding learning. InProceedings of the IEEE international con- ference on computer vision, pages 2840–2848, 2017. 4
2017
-
[43]
Omnicontrol: Control any joint at any time for human motion generation
Yiming Xie, Varun Jampani, Lei Zhong, Deqing Sun, and Huaizu Jiang. Omnicontrol: Control any joint at any time for human motion generation. InInternational Conference on Learning Representations, pages 28176–28194, 2024. 2
2024
-
[44]
Inter-x: Towards versatile human- human interaction analysis
Liang Xu, Xintao Lv, Yichao Yan, Xin Jin, Shuwen Wu, Congsheng Xu, Yifan Liu, Yizhou Zhou, Fengyun Rao, Xingdong Sheng, et al. Inter-x: Towards versatile human- human interaction analysis. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 22260–22271, 2024. 1, 2, 5
2024
-
[45]
Multi-person interaction generation from two-person motion priors
Wenning Xu, Shiyu Fan, Paul Henderson, and Edmond SL Ho. Multi-person interaction generation from two-person motion priors. InProceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Confer- ence Conference Papers, pages 1–11, 2025. 2
2025
-
[46]
Generating human interaction motions in scenes with text control
Hongwei Yi, Justus Thies, Michael J Black, Xue Bin Peng, and Davis Rempe. Generating human interaction motions in scenes with text control. InEuropean Conference on Com- puter Vision, pages 246–263. Springer, 2024. 1
2024
-
[47]
Generating human motion from textual descrip- tions with discrete representations
Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Yong Zhang, Hongwei Zhao, Hongtao Lu, Xi Shen, and Ying Shan. Generating human motion from textual descrip- tions with discrete representations. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14730–14740, 2023. 2
2023
-
[48]
Sgar: Structural generative augmentation for 3d human mo- tion retrieval.Advances in Neural Information Processing Systems, 38:106931–106959, 2026
Jiahang Zhang, Lilang Lin, Shuai Yang, and Jiaying Liu. Sgar: Structural generative augmentation for 3d human mo- tion retrieval.Advances in Neural Information Processing Systems, 38:106931–106959, 2026. 2
2026
-
[49]
Re- modiffuse: Retrieval-augmented motion diffusion model
Mingyuan Zhang, Xinying Guo, Liang Pan, Zhongang Cai, Fangzhou Hong, Huirong Li, Lei Yang, and Ziwei Liu. Re- modiffuse: Retrieval-augmented motion diffusion model. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 364–373, 2023. 2
2023
-
[50]
Ego- body: Human body shape and motion of interacting peo- ple from head-mounted devices
Siwei Zhang, Qianli Ma, Yan Zhang, Zhiyin Qian, Taein Kwon, Marc Pollefeys, Federica Bogo, and Siyu Tang. Ego- body: Human body shape and motion of interacting peo- ple from head-mounted devices. InEuropean conference on computer vision, pages 180–200. Springer, 2022. 1
2022
-
[51]
Zewei Zhang, Kehan Wen, Michael Xu, Junzhe He, Chen- hao Li, Takahiro Miki, Clemens Schwarke, Chong Zhang, Xue Bin Peng, and Marco Hutter. Learning whole-body hu- manoid locomotion via motion generation and motion track- ing.arXiv preprint arXiv:2604.17335, 2026. 1 10 MIME: Multimodal Interactive Motion Encoder Supplementary Material Figure 5. Example of ...
Pith/arXiv arXiv 2026
-
[52]
Shapetalk: A language dataset and framework for 3d shape edits and deforma- tions
Panos Achlioptas, Ian Huang, Minhyuk Sung, Sergey Tulyakov, and Leonidas Guibas. Shapetalk: A language dataset and framework for 3d shape edits and deforma- tions. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 12685– 12694, 2023
2023
-
[53]
Posefix: Correcting 3d hu- man poses with natural language
Ginger Delmas, Philippe Weinzaepfel, Francesc Moreno- Noguer, and Gr ´egory Rogez. Posefix: Correcting 3d hu- man poses with natural language. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 15018–15028, 2023
2023
-
[54]
Yebin Yang, Di Wen, Lei Qi, Weitong Kong, Junwei Zheng, Ruiping Liu, Yufan Chen, Chengzhi Wu, Kailun Yang, Yuqian Fu, et al. Interedit: Navigating text-guided multi- human 3d motion editing.arXiv preprint arXiv:2603.13082, 2026. 14
Pith/arXiv arXiv 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.