REVIEW 5 major objections 6 minor 1 cited by
SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters
T0 review · 5 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read SOLAMI claims to be the first end-to-end social vision-language-action model that takes a user's speech and body motion as input and generates a 3D character's speech and body motion as output, with lower latency than modular LLM-agent…
desk verdict A real end-to-end social VLA system with a useful synthetic data pipeline, but the headline evaluation is weakened by an in-distribution test set and an AI judge drawn from the same model family that generated the data. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is tokenization: user speech is encoded with SpeechTokenizer into semantic tokens and user motion is encoded with three VQ-VAEs for body, hands, and inter-character relative transform, turning both modalities into discrete sequences that the decoder-only LLM can treat as additional languages. The LLM then generates the character's response tokens in a multi-round conversation template marked by special modality tokens. A three-stage training scheme—tokenizer training, multi-task pre-training for motion-text and speech-text alignment, then instruction tuning on SynMSI—is what makes the token spaces usable for social dialogue, and the SynMSI pipeline supplies the interaction data by retrieving motions from text-motion databases and refining LLM-written scripts around them.
What would settle it
Collect a held-out set of real user–character interactions in VR (user speech, tracked body motion, and character responses) and compare SOLAMI with the DLP baseline on motion FID, PA-MPJPE, speech relevance, and end-to-end latency; if SOLAMI's advantage on the synthetic SynMSI test set shrinks or reverses on this real-data evaluation, the central claim that the synthetic-data-trained end-to-end model transfers to live interaction would be contradicted.
Extended reading notes
Core claim
The paper's central claim is that a social VLA model can replace the modular LLM-agent architecture for embodied 3D characters. SOLAMI treats user speech tokens and user motion tokens as inputs to a decoder-only LLM, predicts the character's speech and motion tokens from context and character setting, and decodes them with SoundStorm and VQ-VAE decoders. The authors report that this end-to-end design gives better motion quality (FID, PA-MPJPE, angle error), better context-relevant speech, and lower inference latency than the LLM+Speech and DLP baselines, and that a 60-participant VR user study rates it highest on motion coherence, motion interaction, speech consistency, and overall experience. The paper also claims SynMSI makes training feasible by generating 6.3K multi-turn multimodal conversations from existing motion datasets.
Load-bearing premise
The load-bearing premise is that the synthetic SynMSI dataset—motions retrieved from existing databases with GPT-4o-written scripts and TTS voices—captures real user interaction well enough that training on it transfers to live VR use; if the synthetic distribution does not transfer, the reported gains over the baselines would not reflect real-world performance.
Editorial extensions
If this is right
- A single decoder-only LLM can handle the full loop of understanding user speech and body motion and generating character speech and motion, so the text-based modular pipeline is not required for immersive interaction.
- The end-to-end design reduces inference latency relative to ASR plus LLM plus TTS pipelines, which matters for real-time conversation.
- Pre-training on motion-text and speech-text alignment tasks is necessary; SOLAMI without pretraining scores worse on both motion and speech quality.
- Full-parameter fine-tuning outperforms LoRA fine-tuning for this multimodal dialogue task on the SynMSI data.
- The SynMSI synthesis pipeline can produce large-scale multimodal interaction data from existing motion datasets, which addresses the data scarcity bottleneck for social VLA training.
Reading between the lines
- If the end-to-end design generalizes, the same tokenized speech-plus-motion interface could be extended to vision input (user video) once paired social interaction data exists, giving characters genuine visual perception of the user.
- A direct test of the synthetic-data assumption would be training on SynMSI but evaluating with real captured VR interactions, including noisy tracking, to see whether the motion-quality and latency advantages persist outside the synthetic distribution.
- The pipeline used to build SynMSI could be repurposed to generate multi-party interactions or object-mediated interactions, since it only requires a motion-text database and an LLM script generator.
- The latency advantage is likely to grow as backbones scale, because the end-to-end model avoids accumulating per-module latencies.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents SOLAMI, an end-to-end social vision-language-action (VLA) model for driving 3D autonomous characters in immersive VR interaction. Given the user's speech and body motion as input, the model generates the character's responsive speech and motion tokens through a decoder-only LLM backbone equipped with motion and speech tokenizers. The authors also introduce SynMSI, a synthetic multimodal interaction dataset built from existing motion-text datasets, GPT-4o-generated scripts, and TTS/voice cloning, and they describe a VR interface built on Oculus Quest 3. Quantitative evaluations on a held-out split of SynMSI and a 60-participant VR user study are reported, with the central claim that SOLAMI achieves more precise and natural multimodal responses with lower latency than modular LLM-agent baselines such as DLP and LLM+Speech.
Significance. If the reported results hold, the paper advances an underexplored direction: applying the end-to-end VLA paradigm, which has proven successful in robot manipulation, to social interaction with virtual characters. The main strengths are the coherent unified architecture, the automatic synthetic data pipeline that repurposes existing motion datasets, and the tangible VR testbed for cross-architecture comparison. The paper also gives useful ablations (e.g., full fine-tuning vs. LoRA, pretraining on/off). However, the quantitative evidence for the headline claim is weakened by the evaluation design: the test split is drawn from the same synthetic pipeline that produced the training data, the GPT-4o-based automatic judge belongs to the same model family used to generate the ground-truth scripts, and the reported averages over 5 runs lack variance estimates and statistical tests. The VR user study is the only out-of-distribution evidence, but it is limited to self-reported Likert scores from 60 participants. These issues are fixable in a revision, and the architecture and system contributions remain potentially valuable.
major comments (5)
- [Sec. 6.2, Table 1] The claim of 'significantly outperforms' is not supported by the reported statistics. Table 1 reports only point averages over 5 runs, with no standard deviations, confidence intervals, or significance tests for either the motion metrics or the speech metrics. For example, the difference between SOLAMI (full params) and DLP in FID is 4.254 vs. 3.443, but without variance or a paired test it is impossible to know whether this is meaningful. Please report per-run results or error bars and run appropriate statistical tests (or at least justify that the differences exceed run-to-run variability).
- [Sec. 4.3 and Sec. 6.2] There is a circularity concern in the speech-content evaluation. GPT-4o is used to generate the SynMSI dialogue scripts (Sec. 4.3) and is also used as the judge for Context Relevance and Character Consistency (Sec. 6.2). The near-ceiling scores for the SynMSI Dataset row (4.888 and 4.893) are consistent with this judge-data overlap. I request either a human evaluation with inter-annotator agreement, an independent judge from a different model family, or at least a calibration study showing that GPT-4o as judge agrees with human ratings on a held-out sample.
- [Sec. 6.1 and Sec. 6.2] The baseline comparisons raise fairness questions that should be addressed explicitly. DLP is modified by replacing its MoMat-MoGen module with MotionGPT because the original was 'too slow for user interaction (over 5 seconds latency)', and LLM+Speech truncates responses to at most 3 sentences. The resulting latency comparison (2.639s vs. 5.518s) is therefore between SOLAMI and a modified DLP, not the published DLP system. Please report the latency of the original components, justify the truncation policy's effect on speech-quality scores, and discuss how these changes affect the conclusions.
- [Appendix A and Sec. 6.3] The paper's own future-work section concedes that 'collecting real-time data of actual dyadic interaction could enable our model to generate more precise and natural body language and speech.' This is directly relevant to the synthetic-to-real transfer assumption. Since the quantitative evaluation in Table 1 is on a split of the synthetic SynMSI data, and the only out-of-distribution evidence is a 60-participant self-report study, the paper should temper its generalizability claims or provide additional evidence of transfer (e.g., interaction logs, objective measures of motion understanding in VR, or a larger and more diverse user study).
- [Abstract and Sec. 2] The claim of being 'the first end-to-end Social vision-Language-Action' should be reconciled with prior work cited in the paper itself, especially body-of-her [9], which appears to describe an end-to-end humanoid agent. The related-work section discusses modular LLM agents and motion LLMs but does not position SOLAMI against this prior end-to-end effort. Please clarify the precise technical differences or adjust the 'first' claim.
minor comments (6)
- [Abstract] There is a typo in the second sentence: 'foundamental' should be 'fundamental'. Also, 'an immersive' appears twice in the abstract; please rephrase.
- [Sec. 3.2, Eq. (2)] The hyperparameters lambda_r, lambda_e, lambda_c, and lambda_v are described only as 'manually adjusted weights'. For reproducibility, please provide their concrete values or a link to the configuration.
- [Sec. 3.1] The speech tokenizer is described as SpeechTokenizer, but Sec. 3.2 says the pre-trained checkpoint is taken from AnyGPT. Please clarify the relationship: is the SpeechTokenizer initialized from AnyGPT's checkpoint for the semantic layer, or are they separate components?
- [Sec. 6.3, Table 2] The questionnaire dimension is called 'Speech Consistency' in Table 2 but 'Speech Coherence' in the running text of Sec. 6.3. Please unify the terminology.
- [Table 1] The dashes in Table 1 are explained implicitly, but it would be helpful to add a footnote clarifying that methods without motion output are not evaluated on motion metrics (FID, diversity, PA-MPJPE, angle error), and that the SynMSI Dataset row serves as the reference ground-truth quality.
- [Sec. 5] The VR interface uses Oculus Quest full-body tracking [73], but the description could benefit from a few sentences on how the captured pose data is retargeted to SMPL-X and how the character's motion is retargeted back to the 3D avatar; this is important for reproducibility of the user study.
Circularity Check
Speech-content evaluation is partially circular: GPT-4o writes the SynMSI scripts and then judges Context Relevance and Character Consistency on a test split drawn from that same generated distribution; motion metrics, latency, and the VR user study remain independent.
-
other
[Sec. 4.3, Secs. 6.1-6.2, Table 1]
"Based on the topic, character setting, and previous round of scripts, we use GPT-4o [52] to generate textual descriptions (motion, speech, expression etc.) for the next round of the dialogue. ... We split the synthesized multimodal data into training and test sets with a 9:1 ratio. ... we employ GPT-4o [52] as the judge to assess Context Relevance and Character Consistency on a Likert scale ranging from 1 to 5."
The two speech-content metrics that support the claim of more precise and natural responses are scored by GPT-4o, the same model used to author the SynMSI reference scripts and to brainstorm topics/tasks. The test split is drawn from that same generated distribution, so high Context Relevance and Character Consistency scores partly measure the judge's agreement with its own generation distribution rather than independent conversational quality. Table 1's reference row scores 4.888/4.893 on these metrics, near ceiling, which is consistent with judge-data circularity.
full rationale
The core derivation chain is self-contained: user speech and motion are tokenized, a decoder-only LLM autoregressively predicts character speech and motion tokens, and decoders produce the output modalities. No fitted parameter is renamed as a prediction, and no load-bearing result rests solely on a self-citation. Motion metrics use external feature extractors (AIST++ features for FID/diversity, standard PA-MPJPE and angle error), voice similarity uses UniAudio features, and latency is measured on a deployed vLLM/H800 stack; these are not derived from SynMSI. The VR user study with 60 participants is genuinely out-of-distribution and favors SOLAMI across all dimensions. However, the quantitative speech-content evaluation is partially circular: GPT-4o generates the SynMSI scripts (Sec. 4.3) and also judges Context Relevance and Character Consistency (Sec. 6.2) on a 9:1 split of that same synthetic distribution, with the reference row near ceiling. Appendix A explicitly concedes that collecting real-time data of actual dyadic interaction could enable more precise and natural body language and speech, which further weakens the synthetic-test evidence. Because the central VLA claim retains independent architectural content and external motion/latency/user-study support, the appropriate score is 4 rather than higher.
Assumptions & free parameters
free parameters (2)
- Loss weights lambda_r, lambda_e, lambda_c, lambda_v in Eq. (2) =
not specified
- Speech-motion data sampling ratio 4:6 =
4:6
assumptions (3)
- domain assumption A decoder-only LLM can learn to generate coherent speech and motion tokens from concatenated modality token sequences
- domain assumption SMPL-X joint rotations are a sufficient and retargetable representation for 3D character motion
- ad hoc to paper Synthetic data produced by GPT-4o scripts plus motion retrieval and TTS transfer to real user interactions
Cite this review
Pith. "Pith review of SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters." pith.science (2026). https://pith.science/paper/HQDMT3T2
@misc{pith2026241200174,
author = {Pith},
title = {Pith review of: SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters},
year = {2026},
howpublished = {\url{https://pith.science/paper/HQDMT3T2}},
note = {Machine review of arXiv:2412.00174}
}
read the original abstract
Human beings are social animals. How to equip 3D autonomous characters with similar social intelligence that can perceive, understand and interact with humans remains an open yet foundamental problem. In this paper, we introduce SOLAMI, the first end-to-end Social vision-Language-Action (VLA) Modeling framework for Immersive interaction with 3D autonomous characters. Specifically, SOLAMI builds 3D autonomous characters from three aspects: (1) Social VLA Architecture: We propose a unified social VLA framework to generate multimodal response (speech and motion) based on the user's multimodal input to drive the character for social interaction. (2) Interactive Multimodal Data: We present SynMSI, a synthetic multimodal social interaction dataset generated by an automatic pipeline using only existing motion datasets to address the issue of data scarcity. (3) Immersive VR Interface: We develop a VR interface that enables users to immersively interact with these characters driven by various architectures. Extensive quantitative experiments and user studies demonstrate that our framework leads to more precise and natural character responses (in both speech and motion) that align with user expectations with lower latency.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 1 Pith paper
-
Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis
Human-X jointly predicts actions and reactions in real time to produce physically plausible human-machine interaction motion.
Reference graph
Works this paper leans on
-
[9]
Body of her: A preliminary study on end- to-end humanoid agent
Tenglong Ao. Body of her: A preliminary study on end- to-end humanoid agent. arXiv preprint arXiv:2408.02879 ,
-
[1]
https://elevenlabs.io/
Elevenlabs. https://elevenlabs.io/. 17
-
[2]
https://character.ai
Character.ai. https://character.ai. 2
-
[3]
https : / / trends
Google trends. https : / / trends . google . com / trends. 5, 16
-
[4]
https://www.talkie-ai.com
Talkie ai. https://www.talkie-ai.com. 2
-
[5]
https://vroid.com/en/studio
Vroid studio. https://vroid.com/en/studio. 6
-
[6]
https://www.zhihu.com
Zhihu. https://www.zhihu.com. 5, 16
-
[7]
Chattts: A generative speech model for daily dia- logue
2noise. Chattts: A generative speech model for daily dia- logue. 2024. 17
2024
Show all 100 references
-
[8]
Listen, denoise, action! audio-driven motion synthesis with diffusion models
Simon Alexanderson, Rajmund Nagy, Jonas Beskow, and Gustav Eje Henter. Listen, denoise, action! audio-driven motion synthesis with diffusion models. ACM TOG, 2023. 2
2023
-
[10]
Tyers, and Gregor Weber
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis M. Tyers, and Gregor Weber. Common voice: A massively-multilingual speech corpus. In LREC, 2020. 5
2020
-
[11]
Gen2act: Hu- man video generation in novel scenarios enables generaliz- able robot manipulation
Homanga Bharadhwaj, Debidatta Dwibedi, Abhinav Gupta, Shubham Tulsiani, Carl Doersch, Ted Xiao, Dhruv Shah, Fei Xia, Dorsa Sadigh, and Sean Kirmani. Gen2act: Hu- man video generation in novel scenarios enables generaliz- able robot manipulation. arXiv preprint arXiv:2409.16283,
-
[12]
Audiolm: A language modeling approach to audio generation
Zal ´an Borsos, Rapha ¨el Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matthew Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasac- chi, and Neil Zeghidour. Audiolm: A language modeling approach to audio generation. TASLP, 2023. 3
2023
-
[13]
Soundstorm: Efficient parallel audio generation
Zal ´an Borsos, Matt Sharifi, Damien Vincent, Eugene Kharitonov, Neil Zeghidour, and Marco Tagliasacchi. Soundstorm: Efficient parallel audio generation. arXiv preprint arXiv:2305.09636, 2023. 3, 17
2023 arXiv
-
[14]
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakr- ishnan, Karol Hausman, Alexander Herzog, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Tomas Jackson, Sally Jesmonth, Nikhil J. Joshi, Ryan Julian, Dmitry Kala...
2023
-
[15]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Sub- biah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakan- tan, Pranav Shyam, Girish Sastry, Amanda Askell, Sand- hini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler...
2020
-
[16]
Smpler-x: Scaling up expressive human pose and shape estimation
Zhongang Cai, Wanqi Yin, Ailing Zeng, Chen Wei, Qing- ping Sun, Wang Yanjun, Hui En Pang, Haiyi Mei, Mingyuan Zhang, Lei Zhang, Chen Change Loy, Lei Yang, and Ziwei Liu. Smpler-x: Scaling up expressive human pose and shape estimation. In NeurIPS, 2023. 2, 13, 15
2023
-
[17]
Digital life project: Autonomous 3d characters with social intelligence
Zhongang Cai, Jianping Jiang, Zhongfei Qing, Xinying Guo, Mingyuan Zhang, Zhengyu Lin, Haiyi Mei, Chen Wei, Ruisi Wang, Wanqi Yin, Liang Pan, Xiangyu Fan, Han Du, Peng Gao, Zhitao Yang, Yang Gao, Jiaqi Li, Tianxiang Ren, Yukun Wei, Xiaogang Wang, Chen Change Loy, Lei Yang, and...
2024
-
[18]
Mars5: A novel speech model for insane prosody
CAMB.AI. Mars5: A novel speech model for insane prosody. 2024. 17
2024
-
[19]
Xtts: a mas- sively multilingual zero-shot text-to-speech model
Edresson Casanova, Kelly Davis, Eren G ¨olge, G ¨orkem G¨oknar, Iulian Gulea, Logan Hart, Aya Aljafari, Joshua Meyer, Reuben Morais, Samuel Olayemi, et al. Xtts: a mas- sively multilingual zero-shot text-to-speech model. arXiv preprint arXiv:2406.04904, 2024. 5, 6, 17
2024 arXiv
-
[20]
Gr-2: A gen- erative video-language-action model with web-scale knowl- edge for robot manipulation, 2024
Chi-Lam Cheang, Guangzeng Chen, Ya Jing, Tao Kong, Hang Li, Yifeng Li, Yuxiao Liu, Hongtao Wu, Jiafeng Xu, Yichu Yang, Hanbo Zhang, and Minzhao Zhu. Gr-2: A gen- erative video-language-action model with web-scale knowl- edge for robot manipulation, 2024. 15
2024
-
[21]
Open-television: Teleoperation with immersive active visual feedback
Xuxin Cheng, Jialong Li, Shiqi Yang, Ge Yang, and Xiao- long Wang. Open-television: Teleoperation with immersive active visual feedback. 2024. 13
2024
-
[22]
Agents thinking fast and slow: A talker-reasoner architecture
Konstantina Christakopoulou, Shibl Mourad, and Maja Matari´c. Agents thinking fast and slow: A talker-reasoner architecture. arXiv preprint arXiv:2410.08328, 2024. 13
2024 arXiv
-
[23]
How immer- sive is enough? A meta-analysis of the effect of immersive technology on user presence
James J Cummings and Jeremy N Bailenson. How immer- sive is enough? A meta-analysis of the effect of immersive technology on user presence. Media psychology, 2016. 2
2016
-
[24]
Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duck- worth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussai...
2023
-
[25]
Unitalker: Scaling up audio-driven 3d facial animation through A unified model
Xiangyu Fan, Jiaqi Li, Zhiqian Lin, Weiye Xiao, and Lei Yang. Unitalker: Scaling up audio-driven 3d facial animation through A unified model. 2024. 6
2024
-
[26]
Yao Feng, Jing Lin, Sai Kumar Dwivedi, Yu Sun, Priyanka Patel, and Michael J. Black. Chatpose: Chatting about 3d human pose. In CVPR, 2024. 2
2024
-
[27]
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Mar- tin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Meng...
2022
-
[28]
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zachary Chavis, Joya Chen, Feng Cheng, Fu- Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong, Mar ´...
2024
-
[29]
Generating diverse and natural 3d human motions from text
Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. Generating diverse and natural 3d human motions from text. In CVPR, 2022. 4, 5, 13, 14, 16
2022
-
[30]
Denoising diffu- sion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models. In NeurIPS, 2020. 17
2020
-
[31]
Egolm: Multi-modal language model of egocentric motions
Fangzhou Hong, Vladimir Guzov, Hyo Jin Kim, Yuting Ye, Richard Newcombe, Ziwei Liu, and Lingni Ma. Egolm: Multi-modal language model of egocentric motions. arXiv preprint arXiv:2409.18127, 2024. 2, 3, 6
2024 arXiv
-
[32]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. In ICLR, 2022. 4, 6, 7
2022
-
[33]
An embodied generalist agent in 3d world
Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, and Siyuan Huang. An embodied generalist agent in 3d world. In ICML, 2024. 2, 3
2024
-
[34]
With or without you? Interaction and im- mersion in a virtual reality experience
Sarah Hudson, Sheila Matson-Barkat, Nico Pallamin, and Guillaume Jegou. With or without you? Interaction and im- mersion in a virtual reality experience. Journal of business research, 2019. 2
2019
-
[35]
Joshi, Kyle Jeffrey, Rosario Jauregui Ruano, Jasmine Hsu, Keerthana Gopalakr- ishnan, Byron David, Andy Zeng, and Chuyuan Kelly Fu
Brian Ichter, Anthony Brohan, Yevgen Chebotar, Chelsea Finn, Karol Hausman, Alexander Herzog, Daniel Ho, Ju- lian Ibarz, Alex Irpan, Eric Jang, Ryan Julian, Dmitry Kalashnikov, Sergey Levine, Yao Lu, Carolina Parada, Kan- ishka Rao, Pierre Sermanet, Alexander Toshev, Vincent V...
2022
-
[36]
Human3.6m: Large scale datasets and predic- tive methods for 3d human sensing in natural environments
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6m: Large scale datasets and predic- tive methods for 3d human sensing in natural environments. IEEE TPAMI, 2014. 13
2014
-
[37]
Motiongpt: Human motion as a foreign language
Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu, Gang Yu, and Tao Chen. Motiongpt: Human motion as a foreign language. In NeurIPS, 2023. 2
2023
-
[38]
Motiongpt: Human motion as a foreign language
Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu, Gang Yu, and Tao Chen. Motiongpt: Human motion as a foreign language. In NeurIPS, 2023. 2, 3, 6, 14
2023
-
[39]
Openvla: An open-source vision-language-action model
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Paul Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. Op...
2024 arXiv
-
[40]
Dan Kondratyuk, Lijun Yu, Xiuye Gu, Jos ´e Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vigh- nesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng, Joshua V . Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon...
2024
-
[41]
Gonzalez, Hao 10 Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao 10 Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Pro- ceedings of the ACM SIGOPS 29th Symposium on Operating Sy...
2023
-
[42]
Videochat: Chat-centric video understanding
KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355, 2023. 5, 13
2023 arXiv
-
[43]
Ross, and Angjoo Kanazawa
Ruilong Li, Shan Yang, David A. Ross, and Angjoo Kanazawa. AI choreographer: Music conditioned 3d dance generation with AIST++. In ICCV, 2021. 6
2021
-
[44]
Intergen: Diffusion-based multi-human motion genera- tion under complex interactions
Han Liang, Wenqian Zhang, Wenxuan Li, Jingyi Yu, and Lan Xu. Intergen: Diffusion-based multi-human motion genera- tion under complex interactions. IJCV, 2024. 2
2024
-
[45]
Chathuman: Language-driven 3d human understanding with retrieval-augmented tool reasoning
Jing Lin, Yao Feng, Weiyang Liu, and Michael J Black. Chathuman: Language-driven 3d human understanding with retrieval-augmented tool reasoning. arXiv preprint arXiv:2405.04533, 2024. 2
2024 arXiv
-
[46]
VILA: on pre-training for vi- sual language models
Ji Lin, Hongxu Yin, Wei Ping, Pavlo Molchanov, Moham- mad Shoeybi, and Song Han. VILA: on pre-training for vi- sual language models. In CVPR, 2024. 7
2024
-
[47]
BEAT: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis
Haiyang Liu, Zihao Zhu, Naoya Iwamoto, Yichen Peng, Zhengqing Li, You Zhou, Elif Bozkurt, and Bo Zheng. BEAT: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis. In ECCV,
-
[48]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In NeurIPS, 2023. 5, 15
2023
-
[49]
Humantomato: Text-aligned whole-body motion generation
Shunlin Lu, Ling-Hao Chen, Ailing Zeng, Jing Lin, Ruimao Zhang, Lei Zhang, and Heung-Yeung Shum. Humantomato: Text-aligned whole-body motion generation. InICML, 2024. 3, 14
2024
-
[50]
Text and code embeddings by contrastive pre-training
Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, et al. Text and code embeddings by contrastive pre-training. arXiv preprint arXiv:2201.10005, 2022. 5
2022 arXiv
-
[51]
Bagautdinov, Shao- jie Bai, Trevor Darrell, Angjoo Kanazawa, and Alexander Richard
Evonne Ng, Javier Romero, Timur M. Bagautdinov, Shao- jie Bai, Trevor Darrell, Angjoo Kanazawa, and Alexander Richard. From audio to photoreal embodiment: Synthesiz- ing humans in conversations. In CVPR, 2024. 2, 4
2024
-
[52]
Gpt-4o system card
OpenAI. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024. 2, 3, 4, 5, 6, 15, 16
2024 arXiv
-
[53]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Car- roll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Le...
2022
-
[54]
O’Brien, Carrie Jun Cai, Mered- ith Ringel Morris, Percy Liang, and Michael S
Joon Sung Park, Joseph C. O’Brien, Carrie Jun Cai, Mered- ith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In UIST, 2023. 3
2023
-
[55]
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3d hands, face, and body from a single image. In CVPR, 2019. 14
2019
-
[56]
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3d hands, face, and body from a single image. In CVPR, 2019. 3, 5, 13, 14
2019
-
[57]
Shapellm: Universal 3d object understanding for embodied interaction
Zekun Qi, Runpei Dong, Shaochen Zhang, Haoran Geng, Chunrui Han, Zheng Ge, Li Yi, and Kaisheng Ma. Shapellm: Universal 3d object understanding for embodied interaction. In ECCV, 2024. 13
2024
-
[58]
Openvoice: Versatile instant voice cloning
Zengyi Qin, Wenliang Zhao, Xumin Yu, and Xin Sun. Openvoice: Versatile instant voice cloning. arXiv preprint arXiv:2312.01479, 2023. 17
2023 arXiv
-
[59]
Language models are unsuper- vised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsuper- vised multitask learners. OpenAI blog, 2019. 14
2019
-
[60]
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large-scale weak supervision. In ICML,
-
[61]
Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters. In KDD, 2020. 6
2020
-
[62]
Capideville, Alessandra Preziosa, Francesca Morganti, Daniela Villani, Andrea Gaggioli, Cristina Botella, and Mariano Alca ˜niz Raya
Giuseppe Riva, Fabrizia Mantovani, Claret S. Capideville, Alessandra Preziosa, Francesca Morganti, Daniela Villani, Andrea Gaggioli, Cristina Botella, and Mariano Alca ˜niz Raya. Affective interactions using virtual reality: The link between presence and emotions. Cyberpsychol...
2007
-
[63]
Crowdsourcing virtual reality experiments using vr- chat
David Saffo, Caglar Yildirim, Sara Di Bartolomeo, and Cody Dunne. Crowdsourcing virtual reality experiments using vr- chat. In CHI, 2020. 13
2020
-
[64]
Character-llm: A trainable agent for role-playing
Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. Character-llm: A trainable agent for role-playing. In EMNLP, 2023. 2, 6, 16
2023
-
[65]
Duolando: Follower GPT with off-policy reinforcement learning for dance accompaniment
Li Siyao, Tianpei Gu, Zhitao Yang, Zhengyu Lin, Ziwei Liu, Henghui Ding, Lei Yang, and Chen Change Loy. Duolando: Follower GPT with off-policy reinforcement learning for dance accompaniment. In ICLR, 2024. 2, 6
2024
-
[66]
Inducing illusory ownership of a virtual body
Mel Slater, Daniel P ´erez Marcos, Henrik Ehrsson, and Maria V Sanchez-Vives. Inducing illusory ownership of a virtual body. Frontiers in neuroscience, 2009. 2
2009
-
[67]
Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas L
Theodore R. Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas L. Griffiths. Cognitive architectures for language agents. 2024. 2
2024
-
[68]
Chattts: A generative speech model for daily dia- logue
suno.ai. Chattts: A generative speech model for daily dia- logue. 2023. 17
2023
-
[69]
Llama 2: Open foundation and fine- tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Al- bert, Amjad Almahairi, Yasmine Babaei, Nikolay Bash- lykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhos- ale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fer- nandes, Jere...
2023 arXiv
-
[70]
Charactereval: A chinese benchmark for role-playing conversational agent evaluation
Quan Tu, Shilong Fan, Zihang Tian, Tianhao Shen, Shuo Shang, Xin Gao, and Rui Yan. Charactereval: A chinese benchmark for role-playing conversational agent evaluation. In ACL, 2024. 6
2024
-
[71]
V oyager: An open-ended embodied agent with large language models
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandku- mar. V oyager: An open-ended embodied agent with large language models. TMLR, 2024. 3
2024
-
[72]
Chi, Tat- sunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Bar- ret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tat- sunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. Emergent abilities of large language model...
2022
-
[73]
Winkler, Jungdam Won, and Yuting Ye
Alexander W. Winkler, Jungdam Won, and Yuting Ye. Quest- sim: Human motion tracking from sparse sensors with sim- ulated avatars. In SIGGRAPH Asia, 2022. 5, 15
2022
-
[74]
Unleashing large-scale video generative pre-training for visual robot manipulation
Hongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen, Jiafeng Xu, Xinghang Li, Minghuan Liu, Hang Li, and Tao Kong. Unleashing large-scale video generative pre-training for visual robot manipulation. In ICLR, 2024. 15
2024
-
[75]
Guibas, Dahua Lin, and Gordon Wetzstein
Tong Wu, Guandao Yang, Zhibing Li, Kai Zhang, Ziwei Liu, Leonidas J. Guibas, Dahua Lin, and Gordon Wetzstein. Gpt- 4v(ision) is a human-aligned evaluator for text-to-3d genera- tion. In CVPR, 2024. 13
2024
-
[76]
Inter-x: Towards versatile human-human interaction analysis
Liang Xu, Xintao Lv, Yichao Yan, Xin Jin, Shuwen Wu, Congsheng Xu, Yifan Liu, Yizhou Zhou, Fengyun Rao, Xingdong Sheng, Yunhui Liu, Wenjun Zeng, and Xiaokang Yang. Inter-x: Towards versatile human-human interaction analysis. In CVPR, 2024. 2, 4, 5, 13, 16
2024
-
[77]
Uni- audio: An audio foundation model toward universal audio generation
Dongchao Yang, Jinchuan Tian, Xu Tan, Rongjie Huang, Songxiang Liu, Xuankai Chang, Jiatong Shi, Sheng Zhao, Jiang Bian, Xixin Wu, Zhou Zhao, and Helen Meng. Uni- audio: An audio foundation model toward universal audio generation. arXiv preprint arXiv:2310.00704, 2023. 6
-
[78]
Synbody: Synthetic dataset with layered human models for 3d human perception and modeling
Zhitao Yang, Zhongang Cai, Haiyi Mei, Shuai Liu, Zhaoxi Chen, Weiye Xiao, Yukun Wei, Zhongfei Qing, Chen Wei, Bo Dai, Wayne Wu, Chen Qian, Dahua Lin, Ziwei Liu, and Lei Yang. Synbody: Synthetic dataset with layered human models for 3d human perception and modeling. In ICCV,
-
[79]
Language to rewards for robotic skill synthesis
Wenhao Yu, Nimrod Gileadi, Chuyuan Fu, Sean Kir- mani, Kuang-Huei Lee, Montserrat Gonzalez Arenas, Hao- Tien Lewis Chiang, Tom Erez, Leonard Hasenclever, Jan Humplik, Brian Ichter, Ted Xiao, Peng Xu, Andy Zeng, Tingnan Zhang, Nicolas Heess, Dorsa Sadigh, Jie Tan, Yu- val Tassa...
2023
-
[80]
Soundstream: An end- to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. Soundstream: An end- to-end neural audio codec. TASLP, 2022. 3
2022
-
[81]
Anygpt: Unified multimodal LLM with discrete sequence modeling
Jun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou, Dong Zhang, Zhigeng Liu, Xin Zhang, Ruibin Yuan, Ge Zhang, Linyang Li, Hang Yan, Jie Fu, Tao Gui, Tianxiang Sun, Yu- Gang Jiang, and Xipeng Qiu. Anygpt: Unified multimodal LLM with discrete sequence modeling. In ACL, 2024. 3, 5, 6, 7, 17
2024
-
[82]
Speechgpt: Empow- ering large language models with intrinsic cross-modal con- versational abilities
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. Speechgpt: Empow- ering large language models with intrinsic cross-modal con- versational abilities. In EMNLP, 2023. 3
2023
-
[83]
Motiondif- fuse: Text-driven human motion generation with diffusion model
Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. Motiondif- fuse: Text-driven human motion generation with diffusion model. IEEE TPAMI, 2024. 2
2024
-
[84]
Speechtokenizer: Unified speech tok- enizer for speech large language models
Xin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou, and Xipeng Qiu. Speechtokenizer: Unified speech tok- enizer for speech large language models. arXiv preprint arXiv:2308.16692, 2023. 3, 17
2023 arXiv
-
[85]
Deep long-tailed learning: A survey
Yifan Zhang, Bingyi Kang, Bryan Hooi, Shuicheng Yan, and Jiashi Feng. Deep long-tailed learning: A survey. IEEE TPAMI, 2023. 13
2023
-
[86]
Mo- tiongpt: Finetuned llms are general-purpose motion genera- tors
Yaqi Zhang, Di Huang, Bin Liu, Shixiang Tang, Yan Lu, Lu Chen, Lei Bai, Qi Chu, Nenghai Yu, and Wanli Ouyang. Mo- tiongpt: Finetuned llms are general-purpose motion genera- tors. In AAAI, 2024. 2, 3
2024
-
[87]
Glm-4-voice
Zhipu. Glm-4-voice. 2024. 13
2024
-
[88]
Avatargpt: All- in-one framework for motion understanding, planning, gen- eration and beyond
Zixiang Zhou, Yu Wan, and Baoyuan Wang. Avatargpt: All- in-one framework for motion understanding, planning, gen- eration and beyond. In CVPR, 2024. 2
2024
-
[89]
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mo- hamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. In ICLR, 2024. 5
2024
-
[90]
Tran, Radu Soricut, Anikait Singh, Jaspiar Singh, Pierre Sermanet, Pannag R
Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, Quan Vuong, Vincent Vanhoucke, Huong T. Tran, Radu Soricut, Anikait Singh, Jaspiar Singh, Pierre Sermanet, Pannag R. Sanketi, Grecia Salazar, Michael S. ...
2023
-
[91]
With optimized methods like SMPLify [55], achieving an adequate result requires about 1 second of iteration on a V100 GPU
Time-Consuming Fitting Process: The fitting step is computationally intensive. With optimized methods like SMPLify [55], achieving an adequate result requires about 1 second of iteration on a V100 GPU
-
[92]
bind”), the separate representation (marked as ‘separate
Fitting Artifacts and Distortion: Inevitable fitting er- rors can lead to biologically implausible joint rotations, significantly degrading visual quality. In our experiments, we observed that while human mo- tion representation based on 3D keypoints performs well in terms of ...
-
[93]
Current VR devices’ body tracking systems [73] can- not provide ground truth-level data. For instance, existing VR devices estimate lower body postures instead of captur- ing with wearable sensors, and tracking becomes unreliable when hands move beyond the sensor range of VR e...
-
[94]
Character-related topics: These topics are difficult to collect in bulk from the internet and were generated through GPT-4o [52] brainstorming
-
[95]
News-related topics: Google Trends [3] has compiled many news topics that people care about in daily life
-
[96]
Daily life topics: Some community websites, such as Jike, specifically curate such topic content
-
[97]
(a) Samantha (b) K-VRC (c) Batman (d) Banaya Figure 7
Topics people are curious about: Common Q&A web- sites (such as Quora, Zhihu [6]) specifically organize these topics. (a) Samantha (b) K-VRC (c) Batman (d) Banaya Figure 7. Word cloud visualization of the keywords in the col- lected characters’ topics. After collecting these t...
-
[98]
Method 1: Round-by-Round completion: Using LLM to complete and refine the speech and motion text for each character round by round, which is the method men- tioned in our main paper
-
[99]
Method 2: Character Agent Dialogue: Similar to the So- cioMind approach [17], using two LLMs to play two roles (User and Character), and alternately outputting speech and motion text, followed by refinement
-
[100]
According to our experimental results, Method 1 and Method 2 produce better results
Method 3: One-shot generation: Generating the en- tire multi-turn dialogue script at once, then revising the script round by round based on retrieved motions. According to our experimental results, Method 1 and Method 2 produce better results. Although Method 3 ini- tially gen...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.