REVIEW 1 major objections 1 cited by
Unleashing Infinite Motion: Scaling Expressive Quadrupedal Motion via Generative Video Priors
T0 review · 1 major / 0 minor · reviewed 2026-06-29 · grok-4.3
Pith's one-line read Uni-Mo generates expressive quadruped motions from text prompts using video diffusion models and lifts them into deployable 3D trajectories.
desk verdict The paper's pipeline generates quadruped motions via video diffusion on robot visuals and claims strong real-hardware results, but the 3D lifting step has no reported accuracy checks. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Uni-Mo pipeline that chains LLM prompts, video diffusion synthesis, and Identity Consistency Loss to produce coherent videos that lift reliably into 3D motion references.
What would settle it
If policies trained on the lifted trajectories show deployment success rates well below 90 percent on the physical Unitree Go2, the claim that the generated videos yield usable references would not hold.
Extended reading notes
Core claim
Reframing quadruped motion synthesis as a video generation problem, an LLM proposes motion prompts, a video diffusion model synthesizes the behaviors, and the generated videos are lifted into 3D reference trajectories used to train tracking policies that deploy on physical hardware without any animal data in the loop.
Load-bearing premise
Videos from the diffusion model stay consistent in robot appearance across frames so that accurate 3D trajectories can be extracted and used to train policies that work on real hardware.
Editorial extensions
If this is right
- The released dataset of 7,488 language-annotated motions spanning 18.5 hours supports training of many acrobatic and performative behaviors.
- Tracking policies achieve a 96.7 percent success rate on 392 randomly sampled motions deployed on real hardware.
- A 97.6 percent success rate holds across the full dataset when tested in simulation.
- Expressive motions beyond standard gaits become feasible without animal capture or retargeting steps.
Reading between the lines
- Similar prompt-and-lift pipelines could be adapted to other robot morphologies by changing only the video generation prompts.
- The open dataset may serve as a starting point for combining generated motions with small amounts of real data to improve robustness.
- The same generation approach might extend to tasks involving manipulation or multi-robot coordination once suitable video priors exist.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents Uni-Mo, a fully automated pipeline for scaling expressive quadrupedal motions without animal data. An LLM generates motion prompts, a video diffusion model with a new Identity Consistency Loss synthesizes robot behaviors, the videos are lifted to 3D reference trajectories, and tracking policies are trained and deployed on a real Unitree Go2. The authors release the Quad-Imaginarium dataset (7,488 language-annotated motions) and report 96.7% success on 392 randomly sampled real-robot deployments plus 97.6% success across the full dataset in simulation.
Significance. If the lifting step produces accurate 3D trajectories, the approach would remove a major bottleneck in quadruped motion generation by replacing animal mocap with generative video priors, enabling broader behavioral repertoires. The open release of the 18.5-hour dataset is a concrete community contribution that supports reproducibility and follow-on work.
major comments (1)
- [Abstract and validation experiments] Abstract and validation experiments: the reported 96.7% real-robot and 97.6% simulation success rates are presented without any quantitative metrics on 3D lifting accuracy (e.g., 3D joint position error, reprojection error, foot-skate, or kinematic consistency against mocap or synthetic ground truth). This is load-bearing for the central claim, because policies could achieve high success rates even with depth-ambiguous or artifact-laden references if trained to be robust to noise.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback on the evaluation of the 3D lifting step. We address the major comment below.
read point-by-point responses
-
Referee: [Abstract and validation experiments] Abstract and validation experiments: the reported 96.7% real-robot and 97.6% simulation success rates are presented without any quantitative metrics on 3D lifting accuracy (e.g., 3D joint position error, reprojection error, foot-skate, or kinematic consistency against mocap or synthetic ground truth). This is load-bearing for the central claim, because policies could achieve high success rates even with depth-ambiguous or artifact-laden references if trained to be robust to noise.
Authors: We agree that direct quantitative metrics on 3D lifting accuracy would provide valuable additional evidence and address the concern that high policy success could arise from robustness to noisy references rather than accurate trajectories. The real-robot success rate serves as an end-to-end measure of trajectory usability, but we acknowledge the referee's point that intermediate lifting quality metrics strengthen the central claim. In the revised manuscript we will add such evaluations, including 3D joint position error, reprojection error, foot-skate, and kinematic consistency, computed against synthetic ground truth derived from the video generation process and available kinematic priors. revision: yes
Circularity Check
No significant circularity; derivation relies on external validation
full rationale
The paper describes a pipeline of LLM prompt generation, video diffusion synthesis, 3D lifting via Identity Consistency Loss, policy training, and deployment validation on a real Unitree Go2 robot (96.7% success on 392 samples) plus simulation (97.6%). No equations, fitted parameters, or self-citations are presented that reduce any claimed result to an internal definition or input by construction. The reported success rates are measured on physical hardware and independent simulation benchmarks, making the central claims self-contained against external evaluation rather than tautological. No load-bearing steps match the enumerated circularity patterns.
Assumptions & free parameters
assumptions (2)
- domain assumption Video diffusion models trained on general video data can synthesize coherent quadruped robot motions when conditioned on text prompts
- domain assumption 3D trajectories extracted from generated videos are sufficiently accurate to train deployable tracking policies
Cite this review
Pith. "Pith review of Unleashing Infinite Motion: Scaling Expressive Quadrupedal Motion via Generative Video Priors." pith.science (2026). https://pith.science/paper/B3TXJYMN
@misc{pith2026260628237,
author = {Pith},
title = {Pith review of: Unleashing Infinite Motion: Scaling Expressive Quadrupedal Motion via Generative Video Priors},
year = {2026},
howpublished = {\url{https://pith.science/paper/B3TXJYMN}},
note = {Machine review of arXiv:2606.28237}
}
read the original abstract
Quadruped robots have achieved remarkable locomotion, yet their behavioral repertoire remains confined to a few gaits--far from the expressive, companion-like presence long envisioned for them. Attempts to import the humanoid recipe of large-scale motion data have inherited one tacit assumption: that robot motion must first pass through an animal body, making data collection dependent on cooperative animals, reconstruction fragile across species, and retargeting ill-posed across incompatible morphologies. We propose Uni-Mo, a fully automated pipeline that removes the animal from the loop by reframing data scarcity as a generation problem: an LLM proposes motion prompts, a video diffusion model synthesizes the corresponding robot behaviors, and the generated videos are lifted into 3D reference trajectories used to train tracking policies deployed on a real Unitree Go2. To make naively-drifting generations reliably extractable, we introduce an Identity Consistency Loss that enforces appearance coherence across frames. We release Quad-Imaginarium at https://github.com/Amap-Robotics/Quad-Imaginarium, the resulting open-source dataset of 7,488 language-annotated quadruped motions (18.5 hours) spanning acrobatic and performative behaviors. We validate 392 randomly sampled motions on a real Unitree Go2 with a 96.7% deployment success rate, complemented by a 97.6% success rate across the full dataset in simulation.
Figures
Figures from the paper (6 more)
Forward citations
Cited by 1 Pith paper
-
Behavior Foundations for Quadruped Robots: ABot-C0 Technical Report
A multi-source 16,074-clip quadruped motion library plus a flow-matching generalist tracker shows empirical data scaling and zero-shot unseen tracking, integrated with all-terrain locomotion and real-robot deployment.
Reviewed June 29, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.