REVIEW 3 major objections 5 minor 1 cited by
Grounding Intelligence in Movement
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Movement deserves its own AI foundation model, the paper argues
desk verdict A clear, honest position paper arguing for movement as a first-class AI modality; the agenda is plausible and well-grounded, but the paper leans on a few untested empirical claims and an unresolved tension between cross-species generality and clinical precision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the proposed unified movement modeling framework, which consists of three components: a large aggregated data pile of motion capture, video, wearable sensor, and physiological data across species; a multimodal pretrained backbone that integrates context streams as co-teachers and is trained with biomechanical losses penalizing joint-angle violations, non-physical accelerations, and contact violations; and outcome-aligned benchmarks that reward cross-domain generalization and physical realism. The authors use Moravec's paradox as the motivating observation that motor abilities are easy for animals but hard for AI, and they position movement as a uniquely structured domain because the skeleton is conserved over evolutionary time while the body and world remain deformable.
What would settle it
Train a large movement foundation model on adult human motion data and test it on infant movement or non-human locomotion; if it shows no transfer advantage over a task-specific model trained only on the target domain, the shared-structure assumption fails. A more immediate check would be a formal benchmark of simple interaction generation, such as high-fives or stumbling over a rock, in current video models, since the paper's motivating claim that these systems fail is supported only by informal experiments.
Extended reading notes
Core claim
The central claim is that no current approach, including generative video, pose estimation, action recognition, or world-model-based reinforcement learning, actually models movement as a unified concept. The paper argues that movement has a special structure: it is grounded in embodiment and physics, often admits compact low-dimensional pose representations, and is shared across species because of conserved morphology and physical constraints. A model that learns this structure directly from aggregated cross-modal movement data would generalize across tasks, species, and developmental stages where existing systems fail, such as transferring from adult human locomotion to infant fidgeting or animal gaits. The authors therefore propose that overarching models of movement should be developed as a distinct and essential focus for machine learning.
Load-bearing premise
The whole proposal rests on the premise that movement data from different species, tasks, and sensor types share enough common structure that a single pretrained model can learn a transferable latent space; the paper asserts this but does not test it.
Editorial extensions
If this is right
- Scaling existing video or task-specific models will not by itself produce physical understanding, so a coordinated, purpose-built effort is needed.
- A pretrained movement backbone would give researchers a reusable representation, eliminating duplicated pipelines across medicine, robotics, and conservation.
- Cross-species and cross-modal transfer becomes possible: a model pretrained on human motion could help interpret animal behavior and vice versa.
- Movement-aware losses and physics constraints would make generated motion more realistic and clinically meaningful.
- Benchmarks focused on counterfactual generation and physical perturbation recovery would test causal understanding rather than statistical mimicry.
Reading between the lines
- If the shared-structure premise holds, a movement foundation model should outperform task-specific models on low-shot cross-species transfer, which is a testable prediction the paper does not run.
- The same latent space could link neural activity, muscle signals, and kinematics, potentially unifying motor control and motor decoding research.
- The paper's emphasis on diagnostic micro-signals suggests movement models may need to preserve asymmetries and tremors, which standard augmentations would destroy; this design constraint is implicit in the paper's call for movement-aware augmentations.
- A concrete benchmark of infant-to-adult transfer, or human-to-quadruped transfer, would be the quickest way to test the shared-structure assumption that the framework depends on.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that biological movement should be treated as a primary modeling target for AI, on a par with language and vision. It reviews the fragmented landscape of movement-related models (pose estimation, action recognition, generative motion models, and reinforcement learning / world models), argues that no existing approach models movement as a unified cross-species, cross-modal signal, and proposes a three-step research agenda: aggregate a large multimodal 'movement data pile,' pretrain a biomechanically constrained multimodal backbone, and evaluate on high-impact clinical, ecological, and interactive benchmarks. The paper also discusses ethical risks (privacy, bias, algorithmic phrenology) and considers alternative views, including scaling existing models or relying on task-specific models. The central recommendation is a call to build an overarching movement foundation model.
Significance. If the central claim is correct, the paper identifies a high-value and under-served modeling target, and its proposal to unify data across species, sensors, and contexts could catalyze progress in neuromechanics, clinical assessment, ethology, and human-robot interaction. The manuscript's strengths are that it brings together a broad and current literature, provides a concrete three-step framework rather than only a manifesto, includes a useful Table 1 of high-impact goals, and gives explicit attention to privacy and bias risks. Because it is an argumentative review rather than an empirical study, the usual circular-fit concern does not arise; the relevant evidentiary burden lies in the empirical premises used to motivate the agenda, and several of those premises need support.
major comments (3)
- [Section 3.2, Section 3.3, and Table 1] The feasibility claim that a single pretrained latent space can be both species-invariant enough to support cross-species and developmental transfer (Section 2.3) and precise enough to preserve the millimetric, sub-degree, and millisecond details that the paper itself deems critical for clinical use (Section 2.2, Table 1) is internally in tension. Section 3.2 demands that the backbone 'preserve fine-grained details' and 'strictly respect biomechanical/physical constraints,' while Section 3.3 warns that 'near-miss representations are diagnostically meaningless.' The paper proposes no factorization of the latent space into shared movement primitives and entity- or pathology-specific residuals, nor does it cite evidence that such disentanglement is achievable with current methods; the RL transfer failure in Section 2.3 suggests current models entangle morphology with movement policy. To support the central claim, the authors should either propose a concrete factorization or present a falsifiable benchmark that would test whether a single space can satisfy both requirements.
- [Section 1] The claim that current video generative models 'fail catastrophically at generating simple high-fives or stumbling over a rock' rests on 'informal experiments' with no protocol, model versions, prompts, sample sizes, or evaluation metrics. This claim is load-bearing because it motivates the negative argument that scaling existing video models will not converge on movement understanding; as written it is not verifiable. The authors should either replace it with a reproducible benchmark or clearly mark it as anecdotal motivation for the hypothesis rather than as evidence for it.
- [Section 5 and Section 2.3] The paper asserts that scaling existing approaches will not yield a unified movement model ('scaling alone fails to address core issues...'), but provides no scaling-curve evidence, no counterexample, and no formal argument; the same applies to the assertion in Section 2.3 that an RL agent trained on adult locomotion cannot be adapted to infant fidgeting without full retraining. As a position paper, these are defensible conjectures, but they should be formulated as explicit, falsifiable predictions (e.g., a defined benchmark on which scaled current systems plateau) so that readers can separate the paper's empirical premises from its proposal.
minor comments (5)
- [Section 1] The superscript '1' after 'medicine' has no corresponding footnote in the manuscript; either add the footnote content or remove the superscript.
- [Section 1] The sentence 'All processing of information in the brain, from vision to language, has only one goal: emitting better movement primitives' is an unsupported strong claim; the paper's practical proposal does not depend on this universality, so it should be softened or supported.
- [Section 2.1] The model name 'LLaV A[44]' should be written as 'LLaVA [44]' to match standard usage.
- [Section 3.2] The phrase 'co-teachers (akin to approaches in models like CLIP)' would benefit from a citation to CLIP or a specific variant.
- [Section 2.2] The phrase 'produce very smoothed, average representations' is ambiguous; 'over-smoothed, averaged representations' would be clearer.
Circularity Check
No circularity: an argumentative roadmap whose claims are open empirical hypotheses; no prediction reduces to a fit or a self-citation chain.
full rationale
This is a position paper and research roadmap, not a derivation. It argues that movement should be treated as a primary modeling target and proposes a framework for aggregating movement data, pretraining a multimodal backbone, and evaluating on high-impact use cases. There is no equation that defines a predicted quantity in terms of a fitted parameter, no fitted-input-called-prediction, and no uniqueness theorem or ansatz imported from the authors' prior work. The few self-citations (e.g., prior pose-tracking analysis [73] and neural/time-series foundation-model work [5,6]) are used as contextual evidence for the current fragmentation of the field, not as load-bearing premisses that force the central claim. The central feasibility premise—that movement data across species, tasks, and sensor modalities share enough common structure for a single pretrained latent space—is asserted and remains empirically open, and the paper explicitly concedes that current RL agents do not transfer from adult locomotion to infant fidgeting. An unsupported empirical premise or an internal tension between cross-species generality and diagnosis-grade specificity is a correctness and feasibility risk, not a circularity. No step reduces the conclusion to its own inputs by definition, by construction, or by self-citation, so the circularity score is 0.
Assumptions & free parameters
assumptions (5)
- ad hoc to paper All processing of information in the brain has only one goal: emitting better movement primitives that improve evolutionary fitness.
- domain assumption Movement data across species, tasks, and sensor modalities share enough common structure that a single pretrained latent space can be learned and transfer across domains.
- ad hoc to paper Scaling existing video and task-specific models will not converge on unified movement understanding; purpose-built movement-specific architectures are necessary.
- domain assumption Proper biological movement understanding is possible, as evidenced by animals and humans making inferences from movement, and therefore AI can also achieve it.
- domain assumption Pose-based, compact lower-dimensional representations capture essential movement structure better than raw high-dimensional sensory inputs.
Cite this review
Pith. "Pith review of Grounding Intelligence in Movement." pith.science (2026). https://pith.science/paper/GJORKPOJ
@misc{pith2026250702771,
author = {Pith},
title = {Pith review of: Grounding Intelligence in Movement},
year = {2026},
howpublished = {\url{https://pith.science/paper/GJORKPOJ}},
note = {Machine review of arXiv:2507.02771}
}
read the original abstract
Recent advances in machine learning have dramatically improved our ability to model language, vision, and other high-dimensional data, yet they continue to struggle with one of the most fundamental aspects of biological systems: movement. Across neuroscience, medicine, robotics, and ethology, movement is essential for interpreting behavior, predicting intent, and enabling interaction. Despite its core significance in our intelligence, movement is often treated as an afterthought rather than as a rich and structured modality in its own right. This reflects a deeper fragmentation in how movement data is collected and modeled, often constrained by task-specific goals and domain-specific assumptions. But movement is not domain-bound. It reflects shared physical constraints, conserved morphological structures, and purposeful dynamics that cut across species and settings. We argue that movement should be treated as a primary modeling target for AI. It is inherently structured and grounded in embodiment and physics. This structure, often allowing for compact, lower-dimensional representations (e.g., pose), makes it more interpretable and computationally tractable to model than raw, high-dimensional sensory inputs. Developing models that can learn from and generalize across diverse movement data will not only advance core capabilities in generative modeling and control, but also create a shared foundation for understanding behavior across biological and artificial systems. Movement is not just an outcome, it is a window into how intelligent systems engage with the world.
Figures
Forward citations
Cited by 1 Pith paper
-
Zero-Ablation Overstates Register Content Dependence in DINO Vision Transformers
Zero-ablation overstates register content dependence in DINO ViTs because mean, noise, and cross-image shuffle replacements preserve performance while zeroing does not.
Reference graph
Works this paper leans on
-
[1]
Josh Achiam et al. “Gpt-4 technical report”. In: arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
-
[2]
Cosmos world foundation model platform for physical ai
Niket Agarwal et al. “Cosmos world foundation model platform for physical ai”. In: arXiv preprint arXiv:2501.03575 (2025)
arXiv 2025
-
[3]
Curriculum Reinforcement Learning via Morphology-Environment Co-Evolution
Shuang Ao et al. “Curriculum reinforcement learning via morphology-environment co- evolution”. In: arXiv preprint arXiv:2309.12529 (2023)
work page Pith review arXiv 2023
-
[4]
MNE-BIDS: Organizing electrophysiological data into the BIDS format and facilitating their analysis
Stefan Appelhoff et al. “MNE-BIDS: Organizing electrophysiological data into the BIDS format and facilitating their analysis”. In: Journal of Open Source Software 4.44 (2019), p. 1896
2019
-
[5]
A unified, scalable framework for neural population decoding
Mehdi Azabou et al. “A unified, scalable framework for neural population decoding”. In: Advances in Neural Information Processing Systems 36 (2023), pp. 44937–44956
2023
-
[6]
Relax, it doesn’t matter how you get there: A new self-supervised ap- proach for multi-timescale behavior analysis
Mehdi Azabou et al. “Relax, it doesn’t matter how you get there: A new self-supervised ap- proach for multi-timescale behavior analysis”. In:Advances in Neural Information Processing Systems 36 (2023), pp. 28491–28509
2023
-
[7]
3D bird reconstruction: a dataset, model, and shape recovery from a single view
Marc Badger et al. “3D bird reconstruction: a dataset, model, and shape recovery from a single view”. In: European conference on computer vision. Springer. 2020, pp. 1–17
2020
-
[8]
ChatGarment: Garment Estimation, Generation and Editing via Large Language Models
Siyuan Bian et al. “ChatGarment: Garment Estimation, Generation and Editing via Large Language Models”. In: arXiv preprint arXiv:2412.17811 (2024)
arXiv 2024
Show all 106 references
-
[9]
OmniJet-α: the first cross-task foundation model for particle physics
Joschka Birk, Anna Hallin, and Gregor Kasieczka. “OmniJet-α: the first cross-task foundation model for particle physics”. In: Machine Learning: Science and Technology 5.3 (2024), p. 035031
2024
-
[10]
Video generation models as world simulators
Tim Brooks et al. “Video generation models as world simulators”. In: OpenAI Blog 1 (2024), p. 8
2024
-
[11]
MyoSuite–A contact-rich simulation suite for musculoskeletal motor control
Vittorio Caggiano et al. “MyoSuite–A contact-rich simulation suite for musculoskeletal motor control”. In: arXiv preprint arXiv:2205.13600 (2022)
2022 arXiv
-
[12]
Smpler-x: Scaling up expressive human pose and shape estimation
Zhongang Cai et al. “Smpler-x: Scaling up expressive human pose and shape estimation”. In: Advances in Neural Information Processing Systems 36 (2023), pp. 11454–11468
2023
-
[13]
A short note on the kinetics-700 human action dataset
Joao Carreira et al. “A short note on the kinetics-700 human action dataset”. In:arXiv preprint arXiv:1907.06987 (2019)
2019 arXiv
-
[14]
CAPTURE-24: A large dataset of wrist-worn activity tracker data collected in the wild for human activity recognition
Shing Chan et al. “CAPTURE-24: A large dataset of wrist-worn activity tracker data collected in the wild for human activity recognition”. In: Scientific Data 11.1 (2024), p. 1135
2024
-
[15]
Mammalnet: A large-scale video benchmark for mammal recognition and behavior understanding
Jun Chen et al. “Mammalnet: A large-scale video benchmark for mammal recognition and behavior understanding”. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023, pp. 13052–13061
2023
-
[16]
CoMP: Continual Multimodal Pre-training for Vision Foundation Models
Yitong Chen et al. “CoMP: Continual Multimodal Pre-training for Vision Foundation Models”. In: arXiv preprint arXiv:2503.18931 (2025)
2025 arXiv
-
[17]
Bio-Inspired Motion Emulation for Social Robots: A Real-Time Trajectory Generation and Control Approach
Marvin H Cheng, Po-Lin Huang, and Hao-Chuan Chu. “Bio-Inspired Motion Emulation for Social Robots: A Real-Time Trajectory Generation and Control Approach”. In: Biomimetics 9.9 (2024), p. 557
2024
-
[18]
Muscles in action
Mia Chiquier and Carl V ondrick. “Muscles in action”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023, pp. 22091–22101
2023
-
[19]
The Devonian tetrapod Acanthostega gunnari Jarvik: postcranial anatomy, basal tetrapod interrelationships and patterns of skeletal evolution
MI Coates. “The Devonian tetrapod Acanthostega gunnari Jarvik: postcranial anatomy, basal tetrapod interrelationships and patterns of skeletal evolution”. In: Earth and Environmental Science Transactions of the Royal Society of Edinburgh87.3 (1996), pp. 363–421
1996
-
[20]
Openmmlab 3d human parametric model toolbox and bench- mark
MMHuman3D Contributors. Openmmlab 3d human parametric model toolbox and bench- mark. 2021
2021
-
[21]
Neuro-gpt: Towards a foundation model for eeg
Wenhui Cui et al. “Neuro-gpt: Towards a foundation model for eeg”. In:2024 IEEE Interna- tional Symposium on Biomedical Imaging (ISBI). IEEE. 2024, pp. 1–5
2024
-
[22]
OpenSim: open-source software to create and analyze dynamic sim- ulations of movement
Scott L Delp et al. “OpenSim: open-source software to create and analyze dynamic sim- ulations of movement”. In: IEEE transactions on biomedical engineering 54.11 (2007), pp. 1940–1950
2007
-
[23]
A review of 3D human pose estimation algorithms for markerless motion capture
Yann Desmarais et al. “A review of 3D human pose estimation algorithms for markerless motion capture”. In: Computer Vision and Image Understanding 212 (2021), p. 103275
2021
-
[24]
Muscles of vertebrates: comparative anatomy, evolution, homologies and development
Rui Diogo and Virginia Abdala. Muscles of vertebrates: comparative anatomy, evolution, homologies and development. CRC Press, 2010. 10
2010
-
[25]
WANDR: Intention-guided human motion generation
Markos Diomataris et al. “WANDR: Intention-guided human motion generation”. In:Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, pp. 927–936
2024
-
[26]
Project aria: A new tool for egocentric multi-modal ai research
Jakob Engel et al. “Project aria: A new tool for egocentric multi-modal ai research”. In:arXiv preprint arXiv:2308.13561 (2023)
2023 arXiv
-
[27]
Chatpose: Chatting about 3d human pose
Yao Feng et al. “Chatpose: Chatting about 3d human pose”. In:Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2024, pp. 2093–2103
2024
-
[28]
MoVi: A large multi-purpose human motion and video dataset
Saeed Ghorbani et al. “MoVi: A large multi-purpose human motion and video dataset”. In: Plos one 16.6 (2021), e0253157
2021
-
[29]
Imagebind: One embedding space to bind them all
Rohit Girdhar et al. “Imagebind: One embedding space to bind them all”. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023, pp. 15180– 15190
2023
-
[30]
Moment: A family of open time-series foundation models
Mononito Goswami et al. “Moment: A family of open time-series foundation models”. In: arXiv preprint arXiv:2402.03885 (2024)
2024 arXiv
-
[31]
On memorization in diffusion models
Xiangming Gu et al. “On memorization in diffusion models”. In: arXiv preprint arXiv:2310.02664 (2023)
2023 arXiv
-
[32]
Generating diverse and natural 3d human motions from text
Chuan Guo et al. “Generating diverse and natural 3d human motions from text”. In: Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition . 2022, pp. 5152–5161
2022
-
[33]
World models
David Ha and Jürgen Schmidhuber. “World models”. In:arXiv preprint arXiv:1803.10122 (2018)
2018 arXiv
-
[34]
Mastering diverse domains through world models
Danijar Hafner et al. “Mastering diverse domains through world models”. In: arXiv preprint arXiv:2301.04104 (2023)
2023 arXiv
-
[35]
Motion Diffusion Model for Long Motion Generation
Xiaoxu Han, Hegen Xu, and Huilling Sun. “Motion Diffusion Model for Long Motion Generation”. In: 2024 5th International Conference on Machine Learning and Computer Application (ICMLCA). IEEE. 2024, pp. 603–608
2024
-
[36]
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. “Denoising diffusion probabilistic models”. In: Advances in neural information processing systems 33 (2020), pp. 6840–6851
2020
-
[37]
Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
Catalin Ionescu et al. “Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments”. In: IEEE transactions on pattern analysis and machine intelligence 36.7 (2013), pp. 1325–1339
2013
-
[38]
Motiongpt: Human motion as a foreign language
Biao Jiang et al. “Motiongpt: Human motion as a foreign language”. In: Advances in Neural Information Processing Systems 36 (2023), pp. 20067–20079
2023
-
[39]
Open access dataset and toolbox of high-density surface electromyogram recordings
Xinyu Jiang et al. “Open access dataset and toolbox of high-density surface electromyogram recordings”. In: PhysioNet (2021)
2021
-
[40]
Deep transfer learning in sheep activity recognition using ac- celerometer data
Natasa Kleanthous et al. “Deep transfer learning in sheep activity recognition using ac- celerometer data”. In: Expert Systems with Applications 207 (2022), p. 117925
2022
-
[41]
BioPose: Biomechanically-accurate 3D Pose Estimation from Monoc- ular Videos
Farnoosh Koleini et al. “BioPose: Biomechanically-accurate 3D Pose Estimation from Monoc- ular Videos”. In: arXiv preprint arXiv:2501.07800 (2025)
2025 arXiv
-
[42]
World Model-based Perception for Visual Legged Locomotion
Hang Lai et al. “World Model-based Perception for Visual Legged Locomotion”. In:arXiv preprint arXiv:2409.16784 (2024)
2024 arXiv
-
[43]
Foundation models for time series analysis: A tutorial and survey
Yuxuan Liang et al. “Foundation models for time series analysis: A tutorial and survey”. In: Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining. 2024, pp. 6555–6565
2024
-
[44]
Video-llava: Learning united visual representation by alignment before projection
Bin Lin et al. “Video-llava: Learning united visual representation by alignment before projection”. In: arXiv preprint arXiv:2311.10122 (2023)
2023 arXiv
-
[45]
Motion-x: A large-scale 3d expressive whole-body human motion dataset
Jing Lin et al. “Motion-x: A large-scale 3d expressive whole-body human motion dataset”. In: Advances in Neural Information Processing Systems 36 (2023), pp. 25268–25280
2023
-
[46]
Towards graph foundation models: A survey and beyond
Jiawei Liu et al. “Towards graph foundation models: A survey and beyond”. In:arXiv preprint arXiv:2310.11829 (2023)
2023 arXiv
-
[47]
Unified-io: A unified model for vision, language, and multi-modal tasks
Jiasen Lu et al. “Unified-io: A unified model for vision, language, and multi-modal tasks”. In: arXiv preprint arXiv:2206.08916 (2022)
2022 arXiv
-
[48]
HUMOTO: A 4D Dataset of Mocap Human Object Interactions
Jiaxin Lu et al. “HUMOTO: A 4D Dataset of Mocap Human Object Interactions”. In: arXiv preprint arXiv:2504.10414 (2025). 11
2025
-
[49]
M 3 GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation
Mingshuang Luo et al. “M 3 GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation”. In: arXiv preprint arXiv:2405.16273 (2024)
2024 arXiv
-
[50]
Nymeria: A massive collection of multimodal egocentric daily motion in the wild
Lingni Ma et al. “Nymeria: A massive collection of multimodal egocentric daily motion in the wild”. In: European Conference on Computer Vision. Springer. 2024, pp. 445–465
2024
-
[51]
Chimpact: A longitudinal dataset for understanding chimpanzee be- haviors
Xiaoxuan Ma et al. “Chimpact: A longitudinal dataset for understanding chimpanzee be- haviors”. In: Advances in Neural Information Processing Systems 36 (2023), pp. 27501– 27531
2023
-
[52]
AMASS: Archive of motion capture as surface shapes
Naureen Mahmood et al. “AMASS: Archive of motion capture as surface shapes”. In: Proceedings of the IEEE/CVF international conference on computer vision. 2019, pp. 5442– 5451
2019
-
[53]
FedAAR: A novel federated learning framework for animal activity recogni- tion with wearable sensors
Axiu Mao et al. “FedAAR: A novel federated learning framework for animal activity recogni- tion with wearable sensors”. In: Animals 12.16 (2022), p. 2142
2022
-
[54]
DeepLabCut: markerless pose estimation of user-defined body parts with deep learning
Alexander Mathis et al. “DeepLabCut: markerless pose estimation of user-defined body parts with deep learning”. In: Nature neuroscience 21.9 (2018), pp. 1281–1289
2018
-
[55]
Deep learning, reinforcement learning, and world models
Yutaka Matsuo et al. “Deep learning, reinforcement learning, and world models”. In: Neural Networks 152 (2022), pp. 267–275
2022
-
[56]
Act-ChatGPT: Introducing Action Features into Multi- modal Large Language Models for Video Understanding
Yuto Nakamizo and Keiji Yanai. “Act-ChatGPT: Introducing Action Features into Multi- modal Large Language Models for Video Understanding”. In: International Conference on Pattern Recognition. Springer. 2024, pp. 250–265
2024
-
[57]
Combining Model-based and Data-based approaches for online predic- tions of human trajectories
Aymeric Orhan et al. “Combining Model-based and Data-based approaches for online predic- tions of human trajectories”. In: 2024 10th IEEE RAS/EMBS International Conference for Biomedical Robotics and Biomechatronics (BioRob). IEEE. 2024, pp. 1764–1771
2024
-
[58]
Exploring deep learning techniques for wild animal behaviour clas- sification using animal-borne accelerometers
Ryoma Otsuka et al. “Exploring deep learning techniques for wild animal behaviour clas- sification using animal-borne accelerometers”. In: Methods in Ecology and Evolution 15.4 (2024), pp. 716–731
2024
-
[59]
Real-time EMG based pattern recognition control for hand prosthe- ses: A review on existing methods, challenges and future implementation
Nawadita Parajuli et al. “Real-time EMG based pattern recognition control for hand prosthe- ses: A review on existing methods, challenges and future implementation”. In: Sensors 19.20 (2019), p. 4596
2019
-
[60]
Genie 2: A large-scale foundation world model
J Parker-Holder et al. “Genie 2: A large-scale foundation world model”. In: URL: https://deepmind. google/discover/blog/genie-2-a-large-scale-foundation-world-model (2024)
2024
-
[61]
SLEAP: A deep learning system for multi-animal pose tracking
Talmo D Pereira et al. “SLEAP: A deep learning system for multi-animal pose tracking”. In: Nature methods 19.4 (2022), pp. 486–495
2022
-
[62]
Generating meaning: active inference and the scope and limits of passive AI
Giovanni Pezzulo et al. “Generating meaning: active inference and the scope and limits of passive AI”. In: Trends in Cognitive Sciences 28.2 (2024), pp. 97–112
2024
-
[63]
Wearable sensor-based real-time gait detection: A systematic review
Hari Prasanth et al. “Wearable sensor-based real-time gait detection: A systematic review”. In: Sensors 21.8 (2021), p. 2727
2021
-
[64]
A generic noninvasive neuromotor interface for human- computer interaction
Ctrl-labs at Reality Labs et al. “A generic noninvasive neuromotor interface for human- computer interaction”. In: Biorxiv (2024), pp. 2024–02
2024
-
[65]
A generalist agent
Scott Reed et al. “A generalist agent”. In: arXiv preprint arXiv:2205.06175 (2022)
2022 arXiv
-
[66]
Do robots outperform humans in human-centered domains?
Robert Riener, Luca Rabezzana, and Yves Zimmermann. “Do robots outperform humans in human-centered domains?” In: Frontiers in Robotics and AI 10 (2023), p. 1223946
2023
-
[67]
Embodied hands: Modeling and capturing hands and bodies together
Javier Romero, Dimitrios Tzionas, and Michael J Black. “Embodied hands: Modeling and capturing hands and bodies together”. In: arXiv preprint arXiv:2201.02610 (2022)
2022 arXiv
-
[68]
DiffLocks: Generating 3D Hair from a Single Image using Diffusion Models
Radu Alexandru Rosu et al. “DiffLocks: Generating 3D Hair from a Single Image using Diffusion Models”. In: arXiv preprint arXiv:2505.06166 (2025)
2025 arXiv
-
[69]
Audiopalm: A large language model that can speak and listen
Paul K Rubenstein et al. “Audiopalm: A large language model that can speak and listen”. In: arXiv preprint arXiv:2306.12925 (2023)
2023 arXiv
-
[70]
Human motion trajectory prediction: A survey
Andrey Rudenko et al. “Human motion trajectory prediction: A survey”. In:The International Journal of Robotics Research 39.8 (2020), pp. 895–935
2020
-
[71]
emg2pose: A large and diverse benchmark for surface electromyographic hand pose estimation
Sasha Salter et al. “emg2pose: A large and diverse benchmark for surface electromyographic hand pose estimation”. In: Advances in Neural Information Processing Systems 37 (2024), pp. 55703–55728
2024
-
[72]
Dep-rl: Embodied exploration for reinforcement learning in overac- tuated and musculoskeletal systems
Pierre Schumacher et al. “Dep-rl: Embodied exploration for reinforcement learning in overac- tuated and musculoskeletal systems”. In: arXiv preprint arXiv:2206.00484 (2022). 12
2022 arXiv
-
[73]
Movement science needs different pose tracking algorithms
Nidhi Seethapathi et al. “Movement science needs different pose tracking algorithms”. In: arXiv preprint arXiv:1907.10226 (2019)
2019 arXiv
-
[74]
Reinforcement learning- based motion imitation for physiologically plausible musculoskeletal motor control
Merkourios Simos, Alberto Silvio Chiappa, and Alexander Mathis. “Reinforcement learning- based motion imitation for physiologically plausible musculoskeletal motor control”. In: arXiv preprint arXiv:2503.14637 (2025)
2025
-
[75]
Riemanngfm: Learning a graph foundation model from riemannian geometry
Li Sun et al. “Riemanngfm: Learning a graph foundation model from riemannian geometry”. In: Proceedings of the ACM on Web Conference 2025. 2025, pp. 1154–1165
2025
-
[76]
Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models
Zhiyao Sun et al. “Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models”. In: ACM Transactions on Graphics (TOG) 43.4 (2024), pp. 1–9
2024
-
[77]
Totem: Tokenized time series embed- dings for general time series analysis
Sabera Talukder, Yisong Yue, and Georgia Gkioxari. “Totem: Tokenized time series embed- dings for general time series analysis”. In: arXiv preprint arXiv:2402.16412 (2024)
2024 arXiv
-
[78]
GaitDynamics: A Generative Foundation Model for Analyzing Human Walking and Running
Tian Tan et al. “GaitDynamics: A Generative Foundation Model for Analyzing Human Walking and Running”. In: Research Square (2025), rs–3
2025
-
[79]
Gemini: a family of highly capable multimodal models
Gemini Team et al. “Gemini: a family of highly capable multimodal models”. In: arXiv preprint arXiv:2312.11805 (2023)
2023 arXiv
-
[80]
Human motion diffusion model
Guy Tevet et al. “Human motion diffusion model”. In: arXiv preprint arXiv:2209.14916 (2022)
2022 arXiv
-
[81]
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa. “Mujoco: A physics engine for model-based control”. In: 2012 IEEE/RSJ international conference on intelligent robots and systems. IEEE. 2012, pp. 5026–5033
2012
-
[82]
Llama: Open and efficient foundation language models
Hugo Touvron et al. “Llama: Open and efficient foundation language models”. In: arXiv preprint arXiv:2302.13971 (2023)
2023 arXiv
-
[83]
MammAlps: A multi-view video behavior monitoring dataset of wild mam- mals in the Swiss Alps
Devis Tuia. “MammAlps: A multi-view video behavior monitoring dataset of wild mam- mals in the Swiss Alps”. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025. 2025
2025
-
[84]
Evaluating the world model implicit in a generative model
Keyon Vafa et al. “Evaluating the world model implicit in a generative model”. In:Advances in Neural Information Processing Systems 37 (2024), pp. 26941–26975
2024
-
[85]
MmCows: A Multimodal Dataset for Dairy Cattle Monitoring
Hien Vu et al. “MmCows: A Multimodal Dataset for Dairy Cattle Monitoring”. In: Advances in Neural Information Processing Systems 37 (2024), pp. 59451–59467
2024
-
[86]
Videomae v2: Scaling video masked autoencoders with dual masking
Limin Wang et al. “Videomae v2: Scaling video masked autoencoders with dual masking”. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023, pp. 14549–14560
2023
-
[87]
PromptHMR: Promptable Human Mesh Recovery
Yufu Wang et al. “PromptHMR: Promptable Human Mesh Recovery”. In: arXiv preprint arXiv:2504.06397 (2025)
2025 arXiv
-
[88]
Omnibind: Large-scale omni multimodal representation via binding spaces
Zehan Wang et al. “Omnibind: Large-scale omni multimodal representation via binding spaces”. In: arXiv preprint arXiv:2407.11895 (2024)
2024 arXiv
-
[89]
The application of wearable sensors and machine learning algorithms in rehabilitation training: A systematic review
Suyao Wei and Zhihui Wu. “The application of wearable sensors and machine learning algorithms in rehabilitation training: A systematic review”. In: Sensors 23.18 (2023), p. 7667
2023
-
[90]
Addbiomechanics dataset: Capturing the physics of human motion at scale
Keenon Werling et al. “Addbiomechanics dataset: Capturing the physics of human motion at scale”. In: European Conference on Computer Vision. Springer. 2024, pp. 490–508
2024
-
[91]
ivideogpt: Interactive videogpts are scalable world models
Jialong Wu et al. “ivideogpt: Interactive videogpts are scalable world models”. In: Advances in Neural Information Processing Systems 37 (2024), pp. 68082–68119
2024
-
[92]
Anygraph: Graph foundation model in the wild
Lianghao Xia and Chao Huang. “Anygraph: Graph foundation model in the wild”. In: (2024)
2024
-
[93]
Robot learning in the era of foundation models: A survey
Xuan Xiao et al. “Robot learning in the era of foundation models: A survey”. In: Neurocom- puting (2025), p. 129963
2025
-
[94]
Vitpose: Simple vision transformer baselines for human pose estimation
Yufei Xu et al. “Vitpose: Simple vision transformer baselines for human pose estimation”. In: Advances in neural information processing systems 35 (2022), pp. 38571–38584
2022
-
[95]
Vitpose++: Vision transformer for generic body pose estimation
Yufei Xu et al. “Vitpose++: Vision transformer for generic body pose estimation”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence46.2 (2023), pp. 1212–1230
2023
-
[96]
LLaV Action: evaluating and training multi-modal large language models for action recognition
Shaokai Ye et al. “LLaV Action: evaluating and training multi-modal large language models for action recognition”. In: arXiv preprint arXiv:2503.18712 (2025)
2025
-
[97]
SuperAnimal pretrained pose estimation models for behavioral analysis
Shaokai Ye et al. “SuperAnimal pretrained pose estimation models for behavioral analysis”. In: Nature communications 15.1 (2024), p. 5165. 13
2024
-
[98]
Natural language can help bridge the sim2real gap
Albert Yu et al. “Natural language can help bridge the sim2real gap”. In: arXiv preprint arXiv:2405.10020 (2024)
2024 arXiv
-
[99]
Physdiff: Physics-guided human motion diffusion model
Ye Yuan et al. “Physdiff: Physics-guided human motion diffusion model”. In:Proceedings of the IEEE/CVF international conference on computer vision. 2023, pp. 16010–16021
2023
-
[100]
Motiondiffuse: Text-driven human motion generation with diffusion model
Mingyuan Zhang et al. “Motiondiffuse: Text-driven human motion generation with diffusion model”. In: IEEE transactions on pattern analysis and machine intelligence 46.6 (2024), pp. 4115–4128
2024
-
[101]
Egobody: Human body shape and motion of interacting people from head- mounted devices
Siwei Zhang et al. “Egobody: Human body shape and motion of interacting people from head- mounted devices”. In: European conference on computer vision. Springer. 2022, pp. 180– 200
2022
-
[102]
Meta-transformer: A unified framework for multimodal learning
Yiyuan Zhang et al. “Meta-transformer: A unified framework for multimodal learning”. In: arXiv preprint arXiv:2307.10802 (2023)
2023 arXiv
-
[103]
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
Jiahe Zhao et al. “HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding”. In: arXiv preprint arXiv:2503.12955 (2025)
2025 arXiv
-
[104]
HumanOmni: A Large Vision-Speech Language Model for Human- Centric Video Understanding
Jiaxing Zhao et al. “HumanOmni: A Large Vision-Speech Language Model for Human- Centric Video Understanding”. In: arXiv preprint arXiv:2501.15111 (2025)
2025 arXiv
-
[105]
Detrs beat yolos on real-time object detection
Yian Zhao et al. “Detrs beat yolos on real-time object detection”. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2024, pp. 16965–16974
2024
-
[106]
3D menagerie: Modeling the 3D shape and pose of animals
Silvia Zuffi et al. “3D menagerie: Modeling the 3D shape and pose of animals”. In: Proceed- ings of the IEEE conference on computer vision and pattern recognition. 2017, pp. 6365– 6373. 14 Technical Appendices and Supplementary Material Table 1: Examples of challenging, high-i...
2017
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.