Pith. sign in

REVIEW 1 cited by

CoMo: Controllable Motion Generation through Language Guided Pose Code Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.13900 v2 pith:E6CZWARP submitted 2024-03-20 cs.CV

classification cs.CV
keywords motioncomoeditinggenerationposecodesmodelsmotions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-motion models excel at efficient human motion generation, but existing approaches lack fine-grained controllability over the generation process. Consequently, modifying subtle postures within a motion or inserting new actions at specific moments remains a challenge, limiting the applicability of these methods in diverse scenarios. In light of these challenges, we introduce CoMo, a Controllable Motion generation model, adept at accurately generating and editing motions by leveraging the knowledge priors of large language models (LLMs). Specifically, CoMo decomposes motions into discrete and semantically meaningful pose codes, with each code encapsulating the semantics of a body part, representing elementary information such as "left knee slightly bent". Given textual inputs, CoMo autoregressively generates sequences of pose codes, which are then decoded into 3D motions. Leveraging pose codes as interpretable representations, an LLM can directly intervene in motion editing by adjusting the pose codes according to editing instructions. Experiments demonstrate that CoMo achieves competitive performance in motion generation compared to state-of-the-art models while, in human studies, CoMo substantially surpasses previous work in motion editing abilities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MOST: Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf Interaction

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MOST improves rare-prompt text-to-motion generation by retrieving key motion clips through a new temporal clip Banzhaf interaction and using them as diffusion prompts.

Pith tools