Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Grounding Intelligence in Movement

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Movement deserves its own AI foundation model, the paper argues

desk verdict A clear, honest position paper arguing for movement as a first-class AI modality; the agenda is plausible and well-grounded, but the paper leans on a few untested empirical claims and an unresolved tension between cross-species generality and clinical precision. read the letter →

arxiv 2507.02771 v1 pith:GJORKPOJ submitted 2025-07-03 cs.AI cs.CVcs.LGcs.RO

classification cs.AIcs.CVcs.LGcs.RO
keywords movementfoundationmodelembodiedintelligencecross-speciesgeneralizationbiomechanicalconstraintsmultimodallearningposeestimationactionrecognitionworldmodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This position paper argues that biological movement should be treated as a primary modeling target for artificial intelligence, on par with language and vision, rather than as a byproduct of video, time-series, or task-specific models. The authors contend that movement is a structured, embodied modality governed by shared physical constraints, conserved skeletal geometry, and purposeful dynamics that cut across species and sensors. If they are right, building a purpose-built movement foundation model trained on aggregated data from many species and modalities would be both feasible and superior to simply scaling up existing video generators or reinforcement-learning agents. The paper lays out a three-step roadmap: aggregate a diverse movement data pile, pretrain a multimodal biomechanically constrained backbone, and evaluate on high-impact clinical, ecological, and interactive benchmarks.

What carries the argument

The carrying mechanism is the proposed unified movement modeling framework, which consists of three components: a large aggregated data pile of motion capture, video, wearable sensor, and physiological data across species; a multimodal pretrained backbone that integrates context streams as co-teachers and is trained with biomechanical losses penalizing joint-angle violations, non-physical accelerations, and contact violations; and outcome-aligned benchmarks that reward cross-domain generalization and physical realism. The authors use Moravec's paradox as the motivating observation that motor abilities are easy for animals but hard for AI, and they position movement as a uniquely structured domain because the skeleton is conserved over evolutionary time while the body and world remain deformable.

What would settle it

Train a large movement foundation model on adult human motion data and test it on infant movement or non-human locomotion; if it shows no transfer advantage over a task-specific model trained only on the target domain, the shared-structure assumption fails. A more immediate check would be a formal benchmark of simple interaction generation, such as high-fives or stumbling over a rock, in current video models, since the paper's motivating claim that these systems fail is supported only by informal experiments.

Watch

Extended reading notes

Core claim

The central claim is that no current approach, including generative video, pose estimation, action recognition, or world-model-based reinforcement learning, actually models movement as a unified concept. The paper argues that movement has a special structure: it is grounded in embodiment and physics, often admits compact low-dimensional pose representations, and is shared across species because of conserved morphology and physical constraints. A model that learns this structure directly from aggregated cross-modal movement data would generalize across tasks, species, and developmental stages where existing systems fail, such as transferring from adult human locomotion to infant fidgeting or animal gaits. The authors therefore propose that overarching models of movement should be developed as a distinct and essential focus for machine learning.

Load-bearing premise

The whole proposal rests on the premise that movement data from different species, tasks, and sensor types share enough common structure that a single pretrained model can learn a transferable latent space; the paper asserts this but does not test it.

Editorial extensions

If this is right

  • Scaling existing video or task-specific models will not by itself produce physical understanding, so a coordinated, purpose-built effort is needed.
  • A pretrained movement backbone would give researchers a reusable representation, eliminating duplicated pipelines across medicine, robotics, and conservation.
  • Cross-species and cross-modal transfer becomes possible: a model pretrained on human motion could help interpret animal behavior and vice versa.
  • Movement-aware losses and physics constraints would make generated motion more realistic and clinically meaningful.
  • Benchmarks focused on counterfactual generation and physical perturbation recovery would test causal understanding rather than statistical mimicry.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the shared-structure premise holds, a movement foundation model should outperform task-specific models on low-shot cross-species transfer, which is a testable prediction the paper does not run.
  • The same latent space could link neural activity, muscle signals, and kinematics, potentially unifying motor control and motor decoding research.
  • The paper's emphasis on diagnostic micro-signals suggests movement models may need to preserve asymmetries and tremors, which standard augmentations would destroy; this design constraint is implicit in the paper's call for movement-aware augmentations.
  • A concrete benchmark of infant-to-adult transfer, or human-to-quadruped transfer, would be the quickest way to test the shared-structure assumption that the framework depends on.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper argues that biological movement should be treated as a primary modeling target for AI, on a par with language and vision. It reviews the fragmented landscape of movement-related models (pose estimation, action recognition, generative motion models, and reinforcement learning / world models), argues that no existing approach models movement as a unified cross-species, cross-modal signal, and proposes a three-step research agenda: aggregate a large multimodal 'movement data pile,' pretrain a biomechanically constrained multimodal backbone, and evaluate on high-impact clinical, ecological, and interactive benchmarks. The paper also discusses ethical risks (privacy, bias, algorithmic phrenology) and considers alternative views, including scaling existing models or relying on task-specific models. The central recommendation is a call to build an overarching movement foundation model.

Significance. If the central claim is correct, the paper identifies a high-value and under-served modeling target, and its proposal to unify data across species, sensors, and contexts could catalyze progress in neuromechanics, clinical assessment, ethology, and human-robot interaction. The manuscript's strengths are that it brings together a broad and current literature, provides a concrete three-step framework rather than only a manifesto, includes a useful Table 1 of high-impact goals, and gives explicit attention to privacy and bias risks. Because it is an argumentative review rather than an empirical study, the usual circular-fit concern does not arise; the relevant evidentiary burden lies in the empirical premises used to motivate the agenda, and several of those premises need support.

major comments (3)
  1. [Section 3.2, Section 3.3, and Table 1] The feasibility claim that a single pretrained latent space can be both species-invariant enough to support cross-species and developmental transfer (Section 2.3) and precise enough to preserve the millimetric, sub-degree, and millisecond details that the paper itself deems critical for clinical use (Section 2.2, Table 1) is internally in tension. Section 3.2 demands that the backbone 'preserve fine-grained details' and 'strictly respect biomechanical/physical constraints,' while Section 3.3 warns that 'near-miss representations are diagnostically meaningless.' The paper proposes no factorization of the latent space into shared movement primitives and entity- or pathology-specific residuals, nor does it cite evidence that such disentanglement is achievable with current methods; the RL transfer failure in Section 2.3 suggests current models entangle morphology with movement policy. To support the central claim, the authors should either propose a concrete factorization or present a falsifiable benchmark that would test whether a single space can satisfy both requirements.
  2. [Section 1] The claim that current video generative models 'fail catastrophically at generating simple high-fives or stumbling over a rock' rests on 'informal experiments' with no protocol, model versions, prompts, sample sizes, or evaluation metrics. This claim is load-bearing because it motivates the negative argument that scaling existing video models will not converge on movement understanding; as written it is not verifiable. The authors should either replace it with a reproducible benchmark or clearly mark it as anecdotal motivation for the hypothesis rather than as evidence for it.
  3. [Section 5 and Section 2.3] The paper asserts that scaling existing approaches will not yield a unified movement model ('scaling alone fails to address core issues...'), but provides no scaling-curve evidence, no counterexample, and no formal argument; the same applies to the assertion in Section 2.3 that an RL agent trained on adult locomotion cannot be adapted to infant fidgeting without full retraining. As a position paper, these are defensible conjectures, but they should be formulated as explicit, falsifiable predictions (e.g., a defined benchmark on which scaled current systems plateau) so that readers can separate the paper's empirical premises from its proposal.
minor comments (5)
  1. [Section 1] The superscript '1' after 'medicine' has no corresponding footnote in the manuscript; either add the footnote content or remove the superscript.
  2. [Section 1] The sentence 'All processing of information in the brain, from vision to language, has only one goal: emitting better movement primitives' is an unsupported strong claim; the paper's practical proposal does not depend on this universality, so it should be softened or supported.
  3. [Section 2.1] The model name 'LLaV A[44]' should be written as 'LLaVA [44]' to match standard usage.
  4. [Section 3.2] The phrase 'co-teachers (akin to approaches in models like CLIP)' would benefit from a citation to CLIP or a specific variant.
  5. [Section 2.2] The phrase 'produce very smoothed, average representations' is ambiguous; 'over-smoothed, averaged representations' would be clearer.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: an argumentative roadmap whose claims are open empirical hypotheses; no prediction reduces to a fit or a self-citation chain.

full rationale

This is a position paper and research roadmap, not a derivation. It argues that movement should be treated as a primary modeling target and proposes a framework for aggregating movement data, pretraining a multimodal backbone, and evaluating on high-impact use cases. There is no equation that defines a predicted quantity in terms of a fitted parameter, no fitted-input-called-prediction, and no uniqueness theorem or ansatz imported from the authors' prior work. The few self-citations (e.g., prior pose-tracking analysis [73] and neural/time-series foundation-model work [5,6]) are used as contextual evidence for the current fragmentation of the field, not as load-bearing premisses that force the central claim. The central feasibility premise—that movement data across species, tasks, and sensor modalities share enough common structure for a single pretrained latent space—is asserted and remains empirically open, and the paper explicitly concedes that current RL agents do not transfer from adult locomotion to infant fidgeting. An unsupported empirical premise or an internal tension between cross-species generality and diagnosis-grade specificity is a correctness and feasibility risk, not a circularity. No step reduces the conclusion to its own inputs by definition, by construction, or by self-citation, so the circularity score is 0.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper introduces no free parameters or invented entities. Its central claim rests on several domain assumptions about the structure, universality, and learnability of movement, plus an informal empirical claim about video model failures. The axioms listed above are the load-bearing premises that would need to hold for the proposed unified movement foundation model to succeed.

assumptions (5)
  • ad hoc to paper All processing of information in the brain has only one goal: emitting better movement primitives that improve evolutionary fitness.
    Section 1 states this as a premise, but it is a strong philosophical assertion not standard in neuroscience and not supported by evidence in the paper.
  • domain assumption Movement data across species, tasks, and sensor modalities share enough common structure that a single pretrained latent space can be learned and transfer across domains.
    This is the core feasibility assumption of the proposed unified framework (Abstract, Section 3). It is asserted but never empirically tested.
  • ad hoc to paper Scaling existing video and task-specific models will not converge on unified movement understanding; purpose-built movement-specific architectures are necessary.
    Section 5 (Alternative views) argues this position, but provides no comparative evidence against the scaling alternative.
  • domain assumption Proper biological movement understanding is possible, as evidenced by animals and humans making inferences from movement, and therefore AI can also achieve it.
    Section 1 uses biological performance as existence proof, but does not establish that current AI architectures can exploit the same cues.
  • domain assumption Pose-based, compact lower-dimensional representations capture essential movement structure better than raw high-dimensional sensory inputs.
    The Abstract claims this structure makes movement more interpretable and tractable, but the paper does not compare representation choices empirically.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Grounding Intelligence in Movement." pith.science (2026). https://pith.science/paper/GJORKPOJ

@misc{pith2026250702771,
  author       = {Pith},
  title        = {Pith review of: Grounding Intelligence in Movement},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GJORKPOJ}},
  note         = {Machine review of arXiv:2507.02771}
}
read the original abstract

Recent advances in machine learning have dramatically improved our ability to model language, vision, and other high-dimensional data, yet they continue to struggle with one of the most fundamental aspects of biological systems: movement. Across neuroscience, medicine, robotics, and ethology, movement is essential for interpreting behavior, predicting intent, and enabling interaction. Despite its core significance in our intelligence, movement is often treated as an afterthought rather than as a rich and structured modality in its own right. This reflects a deeper fragmentation in how movement data is collected and modeled, often constrained by task-specific goals and domain-specific assumptions. But movement is not domain-bound. It reflects shared physical constraints, conserved morphological structures, and purposeful dynamics that cut across species and settings. We argue that movement should be treated as a primary modeling target for AI. It is inherently structured and grounded in embodiment and physics. This structure, often allowing for compact, lower-dimensional representations (e.g., pose), makes it more interpretable and computationally tractable to model than raw, high-dimensional sensory inputs. Developing models that can learn from and generalize across diverse movement data will not only advance core capabilities in generative modeling and control, but also create a shared foundation for understanding behavior across biological and artificial systems. Movement is not just an outcome, it is a window into how intelligent systems engage with the world.

Figures

Figures reproduced from arXiv: 2507.02771 by the authors.

Figure 1
Figure 1. Domains with biological movement at their core. Biological movement is central to neuroscience, medicine, computer vision, and sensor modeling—each offering unique but interconnected perspectives on how movement is tracked, modeled, and understood. Biological movement is a truly vexing problem: it lies at a unique intersection between highly structured data and high dimensional unstruc￾tured data. The body itself is… view at source ↗
Figure 2
Figure 2. Unifying framework for movement modeling across species and sensors. The ML community has already developed all of these components. What’s needed is coordinated effort to combine them into a purpose-built framework for learning movement features directly from aggregated movement data. Complementary to RL, world models (e.g., NVIDIA Cosmos[2], Google Genie 2[60]) provide a way for agents to learn a model of their en… view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Zero-Ablation Overstates Register Content Dependence in DINO Vision Transformers

    cs.CV 2026-04 accept novelty 7.0 of 10

    Zero-ablation overstates register content dependence in DINO ViTs because mean, noise, and cross-image shuffle replacements preserve performance while zeroing does not.

Reference graph

Works this paper leans on

106 extracted references · 46 canonical work pages · cited by 1 Pith paper

  1. [1]

    Gpt-4 technical report

    Josh Achiam et al. “Gpt-4 technical report”. In: arXiv preprint arXiv:2303.08774 (2023)

  2. [2]

    Cosmos world foundation model platform for physical ai

    Niket Agarwal et al. “Cosmos world foundation model platform for physical ai”. In: arXiv preprint arXiv:2501.03575 (2025)

  3. [3]

    Curriculum Reinforcement Learning via Morphology-Environment Co-Evolution

    Shuang Ao et al. “Curriculum reinforcement learning via morphology-environment co- evolution”. In: arXiv preprint arXiv:2309.12529 (2023)

  4. [4]

    MNE-BIDS: Organizing electrophysiological data into the BIDS format and facilitating their analysis

    Stefan Appelhoff et al. “MNE-BIDS: Organizing electrophysiological data into the BIDS format and facilitating their analysis”. In: Journal of Open Source Software 4.44 (2019), p. 1896

  5. [5]

    A unified, scalable framework for neural population decoding

    Mehdi Azabou et al. “A unified, scalable framework for neural population decoding”. In: Advances in Neural Information Processing Systems 36 (2023), pp. 44937–44956

  6. [6]

    Relax, it doesn’t matter how you get there: A new self-supervised ap- proach for multi-timescale behavior analysis

    Mehdi Azabou et al. “Relax, it doesn’t matter how you get there: A new self-supervised ap- proach for multi-timescale behavior analysis”. In:Advances in Neural Information Processing Systems 36 (2023), pp. 28491–28509

  7. [7]

    3D bird reconstruction: a dataset, model, and shape recovery from a single view

    Marc Badger et al. “3D bird reconstruction: a dataset, model, and shape recovery from a single view”. In: European conference on computer vision. Springer. 2020, pp. 1–17

  8. [8]

    ChatGarment: Garment Estimation, Generation and Editing via Large Language Models

    Siyuan Bian et al. “ChatGarment: Garment Estimation, Generation and Editing via Large Language Models”. In: arXiv preprint arXiv:2412.17811 (2024)

Show all 106 references
  1. [9]

    OmniJet-α: the first cross-task foundation model for particle physics

    Joschka Birk, Anna Hallin, and Gregor Kasieczka. “OmniJet-α: the first cross-task foundation model for particle physics”. In: Machine Learning: Science and Technology 5.3 (2024), p. 035031

  2. [10]

    Video generation models as world simulators

    Tim Brooks et al. “Video generation models as world simulators”. In: OpenAI Blog 1 (2024), p. 8

  3. [11]

    MyoSuite–A contact-rich simulation suite for musculoskeletal motor control

    Vittorio Caggiano et al. “MyoSuite–A contact-rich simulation suite for musculoskeletal motor control”. In: arXiv preprint arXiv:2205.13600 (2022)

  4. [12]

    Smpler-x: Scaling up expressive human pose and shape estimation

    Zhongang Cai et al. “Smpler-x: Scaling up expressive human pose and shape estimation”. In: Advances in Neural Information Processing Systems 36 (2023), pp. 11454–11468

  5. [13]

    A short note on the kinetics-700 human action dataset

    Joao Carreira et al. “A short note on the kinetics-700 human action dataset”. In:arXiv preprint arXiv:1907.06987 (2019)

  6. [14]

    CAPTURE-24: A large dataset of wrist-worn activity tracker data collected in the wild for human activity recognition

    Shing Chan et al. “CAPTURE-24: A large dataset of wrist-worn activity tracker data collected in the wild for human activity recognition”. In: Scientific Data 11.1 (2024), p. 1135

  7. [15]

    Mammalnet: A large-scale video benchmark for mammal recognition and behavior understanding

    Jun Chen et al. “Mammalnet: A large-scale video benchmark for mammal recognition and behavior understanding”. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023, pp. 13052–13061

  8. [16]

    CoMP: Continual Multimodal Pre-training for Vision Foundation Models

    Yitong Chen et al. “CoMP: Continual Multimodal Pre-training for Vision Foundation Models”. In: arXiv preprint arXiv:2503.18931 (2025)

  9. [17]

    Bio-Inspired Motion Emulation for Social Robots: A Real-Time Trajectory Generation and Control Approach

    Marvin H Cheng, Po-Lin Huang, and Hao-Chuan Chu. “Bio-Inspired Motion Emulation for Social Robots: A Real-Time Trajectory Generation and Control Approach”. In: Biomimetics 9.9 (2024), p. 557

  10. [18]

    Muscles in action

    Mia Chiquier and Carl V ondrick. “Muscles in action”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023, pp. 22091–22101

  11. [19]

    The Devonian tetrapod Acanthostega gunnari Jarvik: postcranial anatomy, basal tetrapod interrelationships and patterns of skeletal evolution

    MI Coates. “The Devonian tetrapod Acanthostega gunnari Jarvik: postcranial anatomy, basal tetrapod interrelationships and patterns of skeletal evolution”. In: Earth and Environmental Science Transactions of the Royal Society of Edinburgh87.3 (1996), pp. 363–421

  12. [20]

    Openmmlab 3d human parametric model toolbox and bench- mark

    MMHuman3D Contributors. Openmmlab 3d human parametric model toolbox and bench- mark. 2021

  13. [21]

    Neuro-gpt: Towards a foundation model for eeg

    Wenhui Cui et al. “Neuro-gpt: Towards a foundation model for eeg”. In:2024 IEEE Interna- tional Symposium on Biomedical Imaging (ISBI). IEEE. 2024, pp. 1–5

  14. [22]

    OpenSim: open-source software to create and analyze dynamic sim- ulations of movement

    Scott L Delp et al. “OpenSim: open-source software to create and analyze dynamic sim- ulations of movement”. In: IEEE transactions on biomedical engineering 54.11 (2007), pp. 1940–1950

  15. [23]

    A review of 3D human pose estimation algorithms for markerless motion capture

    Yann Desmarais et al. “A review of 3D human pose estimation algorithms for markerless motion capture”. In: Computer Vision and Image Understanding 212 (2021), p. 103275

  16. [24]

    Muscles of vertebrates: comparative anatomy, evolution, homologies and development

    Rui Diogo and Virginia Abdala. Muscles of vertebrates: comparative anatomy, evolution, homologies and development. CRC Press, 2010. 10

  17. [25]

    WANDR: Intention-guided human motion generation

    Markos Diomataris et al. “WANDR: Intention-guided human motion generation”. In:Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, pp. 927–936

  18. [26]

    Project aria: A new tool for egocentric multi-modal ai research

    Jakob Engel et al. “Project aria: A new tool for egocentric multi-modal ai research”. In:arXiv preprint arXiv:2308.13561 (2023)

  19. [27]

    Chatpose: Chatting about 3d human pose

    Yao Feng et al. “Chatpose: Chatting about 3d human pose”. In:Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2024, pp. 2093–2103

  20. [28]

    MoVi: A large multi-purpose human motion and video dataset

    Saeed Ghorbani et al. “MoVi: A large multi-purpose human motion and video dataset”. In: Plos one 16.6 (2021), e0253157

  21. [29]

    Imagebind: One embedding space to bind them all

    Rohit Girdhar et al. “Imagebind: One embedding space to bind them all”. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023, pp. 15180– 15190

  22. [30]

    Moment: A family of open time-series foundation models

    Mononito Goswami et al. “Moment: A family of open time-series foundation models”. In: arXiv preprint arXiv:2402.03885 (2024)

  23. [31]

    On memorization in diffusion models

    Xiangming Gu et al. “On memorization in diffusion models”. In: arXiv preprint arXiv:2310.02664 (2023)

  24. [32]

    Generating diverse and natural 3d human motions from text

    Chuan Guo et al. “Generating diverse and natural 3d human motions from text”. In: Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition . 2022, pp. 5152–5161

  25. [33]

    World models

    David Ha and Jürgen Schmidhuber. “World models”. In:arXiv preprint arXiv:1803.10122 (2018)

  26. [34]

    Mastering diverse domains through world models

    Danijar Hafner et al. “Mastering diverse domains through world models”. In: arXiv preprint arXiv:2301.04104 (2023)

  27. [35]

    Motion Diffusion Model for Long Motion Generation

    Xiaoxu Han, Hegen Xu, and Huilling Sun. “Motion Diffusion Model for Long Motion Generation”. In: 2024 5th International Conference on Machine Learning and Computer Application (ICMLCA). IEEE. 2024, pp. 603–608

  28. [36]

    Denoising diffusion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. “Denoising diffusion probabilistic models”. In: Advances in neural information processing systems 33 (2020), pp. 6840–6851

  29. [37]

    Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments

    Catalin Ionescu et al. “Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments”. In: IEEE transactions on pattern analysis and machine intelligence 36.7 (2013), pp. 1325–1339

  30. [38]

    Motiongpt: Human motion as a foreign language

    Biao Jiang et al. “Motiongpt: Human motion as a foreign language”. In: Advances in Neural Information Processing Systems 36 (2023), pp. 20067–20079

  31. [39]

    Open access dataset and toolbox of high-density surface electromyogram recordings

    Xinyu Jiang et al. “Open access dataset and toolbox of high-density surface electromyogram recordings”. In: PhysioNet (2021)

  32. [40]

    Deep transfer learning in sheep activity recognition using ac- celerometer data

    Natasa Kleanthous et al. “Deep transfer learning in sheep activity recognition using ac- celerometer data”. In: Expert Systems with Applications 207 (2022), p. 117925

  33. [41]

    BioPose: Biomechanically-accurate 3D Pose Estimation from Monoc- ular Videos

    Farnoosh Koleini et al. “BioPose: Biomechanically-accurate 3D Pose Estimation from Monoc- ular Videos”. In: arXiv preprint arXiv:2501.07800 (2025)

  34. [42]

    World Model-based Perception for Visual Legged Locomotion

    Hang Lai et al. “World Model-based Perception for Visual Legged Locomotion”. In:arXiv preprint arXiv:2409.16784 (2024)

  35. [43]

    Foundation models for time series analysis: A tutorial and survey

    Yuxuan Liang et al. “Foundation models for time series analysis: A tutorial and survey”. In: Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining. 2024, pp. 6555–6565

  36. [44]

    Video-llava: Learning united visual representation by alignment before projection

    Bin Lin et al. “Video-llava: Learning united visual representation by alignment before projection”. In: arXiv preprint arXiv:2311.10122 (2023)

  37. [45]

    Motion-x: A large-scale 3d expressive whole-body human motion dataset

    Jing Lin et al. “Motion-x: A large-scale 3d expressive whole-body human motion dataset”. In: Advances in Neural Information Processing Systems 36 (2023), pp. 25268–25280

  38. [46]

    Towards graph foundation models: A survey and beyond

    Jiawei Liu et al. “Towards graph foundation models: A survey and beyond”. In:arXiv preprint arXiv:2310.11829 (2023)

  39. [47]

    Unified-io: A unified model for vision, language, and multi-modal tasks

    Jiasen Lu et al. “Unified-io: A unified model for vision, language, and multi-modal tasks”. In: arXiv preprint arXiv:2206.08916 (2022)

  40. [48]

    HUMOTO: A 4D Dataset of Mocap Human Object Interactions

    Jiaxin Lu et al. “HUMOTO: A 4D Dataset of Mocap Human Object Interactions”. In: arXiv preprint arXiv:2504.10414 (2025). 11

  41. [49]

    M 3 GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation

    Mingshuang Luo et al. “M 3 GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation”. In: arXiv preprint arXiv:2405.16273 (2024)

  42. [50]

    Nymeria: A massive collection of multimodal egocentric daily motion in the wild

    Lingni Ma et al. “Nymeria: A massive collection of multimodal egocentric daily motion in the wild”. In: European Conference on Computer Vision. Springer. 2024, pp. 445–465

  43. [51]

    Chimpact: A longitudinal dataset for understanding chimpanzee be- haviors

    Xiaoxuan Ma et al. “Chimpact: A longitudinal dataset for understanding chimpanzee be- haviors”. In: Advances in Neural Information Processing Systems 36 (2023), pp. 27501– 27531

  44. [52]

    AMASS: Archive of motion capture as surface shapes

    Naureen Mahmood et al. “AMASS: Archive of motion capture as surface shapes”. In: Proceedings of the IEEE/CVF international conference on computer vision. 2019, pp. 5442– 5451

  45. [53]

    FedAAR: A novel federated learning framework for animal activity recogni- tion with wearable sensors

    Axiu Mao et al. “FedAAR: A novel federated learning framework for animal activity recogni- tion with wearable sensors”. In: Animals 12.16 (2022), p. 2142

  46. [54]

    DeepLabCut: markerless pose estimation of user-defined body parts with deep learning

    Alexander Mathis et al. “DeepLabCut: markerless pose estimation of user-defined body parts with deep learning”. In: Nature neuroscience 21.9 (2018), pp. 1281–1289

  47. [55]

    Deep learning, reinforcement learning, and world models

    Yutaka Matsuo et al. “Deep learning, reinforcement learning, and world models”. In: Neural Networks 152 (2022), pp. 267–275

  48. [56]

    Act-ChatGPT: Introducing Action Features into Multi- modal Large Language Models for Video Understanding

    Yuto Nakamizo and Keiji Yanai. “Act-ChatGPT: Introducing Action Features into Multi- modal Large Language Models for Video Understanding”. In: International Conference on Pattern Recognition. Springer. 2024, pp. 250–265

  49. [57]

    Combining Model-based and Data-based approaches for online predic- tions of human trajectories

    Aymeric Orhan et al. “Combining Model-based and Data-based approaches for online predic- tions of human trajectories”. In: 2024 10th IEEE RAS/EMBS International Conference for Biomedical Robotics and Biomechatronics (BioRob). IEEE. 2024, pp. 1764–1771

  50. [58]

    Exploring deep learning techniques for wild animal behaviour clas- sification using animal-borne accelerometers

    Ryoma Otsuka et al. “Exploring deep learning techniques for wild animal behaviour clas- sification using animal-borne accelerometers”. In: Methods in Ecology and Evolution 15.4 (2024), pp. 716–731

  51. [59]

    Real-time EMG based pattern recognition control for hand prosthe- ses: A review on existing methods, challenges and future implementation

    Nawadita Parajuli et al. “Real-time EMG based pattern recognition control for hand prosthe- ses: A review on existing methods, challenges and future implementation”. In: Sensors 19.20 (2019), p. 4596

  52. [60]

    Genie 2: A large-scale foundation world model

    J Parker-Holder et al. “Genie 2: A large-scale foundation world model”. In: URL: https://deepmind. google/discover/blog/genie-2-a-large-scale-foundation-world-model (2024)

  53. [61]

    SLEAP: A deep learning system for multi-animal pose tracking

    Talmo D Pereira et al. “SLEAP: A deep learning system for multi-animal pose tracking”. In: Nature methods 19.4 (2022), pp. 486–495

  54. [62]

    Generating meaning: active inference and the scope and limits of passive AI

    Giovanni Pezzulo et al. “Generating meaning: active inference and the scope and limits of passive AI”. In: Trends in Cognitive Sciences 28.2 (2024), pp. 97–112

  55. [63]

    Wearable sensor-based real-time gait detection: A systematic review

    Hari Prasanth et al. “Wearable sensor-based real-time gait detection: A systematic review”. In: Sensors 21.8 (2021), p. 2727

  56. [64]

    A generic noninvasive neuromotor interface for human- computer interaction

    Ctrl-labs at Reality Labs et al. “A generic noninvasive neuromotor interface for human- computer interaction”. In: Biorxiv (2024), pp. 2024–02

  57. [65]

    A generalist agent

    Scott Reed et al. “A generalist agent”. In: arXiv preprint arXiv:2205.06175 (2022)

  58. [66]

    Do robots outperform humans in human-centered domains?

    Robert Riener, Luca Rabezzana, and Yves Zimmermann. “Do robots outperform humans in human-centered domains?” In: Frontiers in Robotics and AI 10 (2023), p. 1223946

  59. [67]

    Embodied hands: Modeling and capturing hands and bodies together

    Javier Romero, Dimitrios Tzionas, and Michael J Black. “Embodied hands: Modeling and capturing hands and bodies together”. In: arXiv preprint arXiv:2201.02610 (2022)

  60. [68]

    DiffLocks: Generating 3D Hair from a Single Image using Diffusion Models

    Radu Alexandru Rosu et al. “DiffLocks: Generating 3D Hair from a Single Image using Diffusion Models”. In: arXiv preprint arXiv:2505.06166 (2025)

  61. [69]

    Audiopalm: A large language model that can speak and listen

    Paul K Rubenstein et al. “Audiopalm: A large language model that can speak and listen”. In: arXiv preprint arXiv:2306.12925 (2023)

  62. [70]

    Human motion trajectory prediction: A survey

    Andrey Rudenko et al. “Human motion trajectory prediction: A survey”. In:The International Journal of Robotics Research 39.8 (2020), pp. 895–935

  63. [71]

    emg2pose: A large and diverse benchmark for surface electromyographic hand pose estimation

    Sasha Salter et al. “emg2pose: A large and diverse benchmark for surface electromyographic hand pose estimation”. In: Advances in Neural Information Processing Systems 37 (2024), pp. 55703–55728

  64. [72]

    Dep-rl: Embodied exploration for reinforcement learning in overac- tuated and musculoskeletal systems

    Pierre Schumacher et al. “Dep-rl: Embodied exploration for reinforcement learning in overac- tuated and musculoskeletal systems”. In: arXiv preprint arXiv:2206.00484 (2022). 12

  65. [73]

    Movement science needs different pose tracking algorithms

    Nidhi Seethapathi et al. “Movement science needs different pose tracking algorithms”. In: arXiv preprint arXiv:1907.10226 (2019)

  66. [74]

    Reinforcement learning- based motion imitation for physiologically plausible musculoskeletal motor control

    Merkourios Simos, Alberto Silvio Chiappa, and Alexander Mathis. “Reinforcement learning- based motion imitation for physiologically plausible musculoskeletal motor control”. In: arXiv preprint arXiv:2503.14637 (2025)

  67. [75]

    Riemanngfm: Learning a graph foundation model from riemannian geometry

    Li Sun et al. “Riemanngfm: Learning a graph foundation model from riemannian geometry”. In: Proceedings of the ACM on Web Conference 2025. 2025, pp. 1154–1165

  68. [76]

    Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models

    Zhiyao Sun et al. “Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models”. In: ACM Transactions on Graphics (TOG) 43.4 (2024), pp. 1–9

  69. [77]

    Totem: Tokenized time series embed- dings for general time series analysis

    Sabera Talukder, Yisong Yue, and Georgia Gkioxari. “Totem: Tokenized time series embed- dings for general time series analysis”. In: arXiv preprint arXiv:2402.16412 (2024)

  70. [78]

    GaitDynamics: A Generative Foundation Model for Analyzing Human Walking and Running

    Tian Tan et al. “GaitDynamics: A Generative Foundation Model for Analyzing Human Walking and Running”. In: Research Square (2025), rs–3

  71. [79]

    Gemini: a family of highly capable multimodal models

    Gemini Team et al. “Gemini: a family of highly capable multimodal models”. In: arXiv preprint arXiv:2312.11805 (2023)

  72. [80]

    Human motion diffusion model

    Guy Tevet et al. “Human motion diffusion model”. In: arXiv preprint arXiv:2209.14916 (2022)

  73. [81]

    Mujoco: A physics engine for model-based control

    Emanuel Todorov, Tom Erez, and Yuval Tassa. “Mujoco: A physics engine for model-based control”. In: 2012 IEEE/RSJ international conference on intelligent robots and systems. IEEE. 2012, pp. 5026–5033

  74. [82]

    Llama: Open and efficient foundation language models

    Hugo Touvron et al. “Llama: Open and efficient foundation language models”. In: arXiv preprint arXiv:2302.13971 (2023)

  75. [83]

    MammAlps: A multi-view video behavior monitoring dataset of wild mam- mals in the Swiss Alps

    Devis Tuia. “MammAlps: A multi-view video behavior monitoring dataset of wild mam- mals in the Swiss Alps”. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025. 2025

  76. [84]

    Evaluating the world model implicit in a generative model

    Keyon Vafa et al. “Evaluating the world model implicit in a generative model”. In:Advances in Neural Information Processing Systems 37 (2024), pp. 26941–26975

  77. [85]

    MmCows: A Multimodal Dataset for Dairy Cattle Monitoring

    Hien Vu et al. “MmCows: A Multimodal Dataset for Dairy Cattle Monitoring”. In: Advances in Neural Information Processing Systems 37 (2024), pp. 59451–59467

  78. [86]

    Videomae v2: Scaling video masked autoencoders with dual masking

    Limin Wang et al. “Videomae v2: Scaling video masked autoencoders with dual masking”. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023, pp. 14549–14560

  79. [87]

    PromptHMR: Promptable Human Mesh Recovery

    Yufu Wang et al. “PromptHMR: Promptable Human Mesh Recovery”. In: arXiv preprint arXiv:2504.06397 (2025)

  80. [88]

    Omnibind: Large-scale omni multimodal representation via binding spaces

    Zehan Wang et al. “Omnibind: Large-scale omni multimodal representation via binding spaces”. In: arXiv preprint arXiv:2407.11895 (2024)

  81. [89]

    The application of wearable sensors and machine learning algorithms in rehabilitation training: A systematic review

    Suyao Wei and Zhihui Wu. “The application of wearable sensors and machine learning algorithms in rehabilitation training: A systematic review”. In: Sensors 23.18 (2023), p. 7667

  82. [90]

    Addbiomechanics dataset: Capturing the physics of human motion at scale

    Keenon Werling et al. “Addbiomechanics dataset: Capturing the physics of human motion at scale”. In: European Conference on Computer Vision. Springer. 2024, pp. 490–508

  83. [91]

    ivideogpt: Interactive videogpts are scalable world models

    Jialong Wu et al. “ivideogpt: Interactive videogpts are scalable world models”. In: Advances in Neural Information Processing Systems 37 (2024), pp. 68082–68119

  84. [92]

    Anygraph: Graph foundation model in the wild

    Lianghao Xia and Chao Huang. “Anygraph: Graph foundation model in the wild”. In: (2024)

  85. [93]

    Robot learning in the era of foundation models: A survey

    Xuan Xiao et al. “Robot learning in the era of foundation models: A survey”. In: Neurocom- puting (2025), p. 129963

  86. [94]

    Vitpose: Simple vision transformer baselines for human pose estimation

    Yufei Xu et al. “Vitpose: Simple vision transformer baselines for human pose estimation”. In: Advances in neural information processing systems 35 (2022), pp. 38571–38584

  87. [95]

    Vitpose++: Vision transformer for generic body pose estimation

    Yufei Xu et al. “Vitpose++: Vision transformer for generic body pose estimation”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence46.2 (2023), pp. 1212–1230

  88. [96]

    LLaV Action: evaluating and training multi-modal large language models for action recognition

    Shaokai Ye et al. “LLaV Action: evaluating and training multi-modal large language models for action recognition”. In: arXiv preprint arXiv:2503.18712 (2025)

  89. [97]

    SuperAnimal pretrained pose estimation models for behavioral analysis

    Shaokai Ye et al. “SuperAnimal pretrained pose estimation models for behavioral analysis”. In: Nature communications 15.1 (2024), p. 5165. 13

  90. [98]

    Natural language can help bridge the sim2real gap

    Albert Yu et al. “Natural language can help bridge the sim2real gap”. In: arXiv preprint arXiv:2405.10020 (2024)

  91. [99]

    Physdiff: Physics-guided human motion diffusion model

    Ye Yuan et al. “Physdiff: Physics-guided human motion diffusion model”. In:Proceedings of the IEEE/CVF international conference on computer vision. 2023, pp. 16010–16021

  92. [100]

    Motiondiffuse: Text-driven human motion generation with diffusion model

    Mingyuan Zhang et al. “Motiondiffuse: Text-driven human motion generation with diffusion model”. In: IEEE transactions on pattern analysis and machine intelligence 46.6 (2024), pp. 4115–4128

  93. [101]

    Egobody: Human body shape and motion of interacting people from head- mounted devices

    Siwei Zhang et al. “Egobody: Human body shape and motion of interacting people from head- mounted devices”. In: European conference on computer vision. Springer. 2022, pp. 180– 200

  94. [102]

    Meta-transformer: A unified framework for multimodal learning

    Yiyuan Zhang et al. “Meta-transformer: A unified framework for multimodal learning”. In: arXiv preprint arXiv:2307.10802 (2023)

  95. [103]

    HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding

    Jiahe Zhao et al. “HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding”. In: arXiv preprint arXiv:2503.12955 (2025)

  96. [104]

    HumanOmni: A Large Vision-Speech Language Model for Human- Centric Video Understanding

    Jiaxing Zhao et al. “HumanOmni: A Large Vision-Speech Language Model for Human- Centric Video Understanding”. In: arXiv preprint arXiv:2501.15111 (2025)

  97. [105]

    Detrs beat yolos on real-time object detection

    Yian Zhao et al. “Detrs beat yolos on real-time object detection”. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2024, pp. 16965–16974

  98. [106]

    3D menagerie: Modeling the 3D shape and pose of animals

    Silvia Zuffi et al. “3D menagerie: Modeling the 3D shape and pose of animals”. In: Proceed- ings of the IEEE conference on computer vision and pattern recognition. 2017, pp. 6365– 6373. 14 Technical Appendices and Supplementary Material Table 1: Examples of challenging, high-i...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.