Pith. sign in

Paper Citation Record · LEDGER

CoMA: Compositional Human Motion Generation with Multi-modal Agents

As of 23 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 5 inbound Pith citation observations for arXiv:2412.07320.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.07320 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:00:10.571415Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:19:20.685328Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T11:51:20.243250Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy34
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 61c33c22-48d6-46cb-a7d1-a16b9a71e2da · outbound

This paper cites GPT-4 Technical Report.

CoMA: Compositional Human Motion Generation with Multi-modal Agents GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:00:10.372795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:00:10.372795Z digest=sha256:6e8faf944b24d4e37691660367c8b84fa1adfa46da23eee595bf50c9049b62d9

Observation edfc546a-ef57-4ca7-a15f-6bd52badd79e · outbound

This paper cites Lan- guage2pose: Natural language grounded pose forecasting.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Lan- guage2pose: Natural language grounded pose forecasting

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:11.258921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.378164Z digest=sha256:43b2169cc44d4391c81a780f47cd66bc47fb372ec1426b046ec2a73a2539a72c

Observation d8cfe53a-c757-4f2d-9681-ce2c0d44af8c · outbound

This paper cites Teach: Temporal action composition for 3d hu- mans.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Teach: Temporal action composition for 3d hu- mans

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:11.244612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.382326Z digest=sha256:aa7ec8e89140af29359b11e4322116b1b28586aeab0387b47a6f40e6c9d6d574

Observation c148009a-93c8-436d-a8bf-c6ce7ae2830f · outbound

This paper cites Black, and Gül Varol.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Black, and Gül Varol

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:11.230256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.386391Z digest=sha256:387bc3bcb7034badb93b717518c4fa3bbddc25432149f84f8275d7f5279a2838

Observation 39c2907e-97ea-4dcd-a6af-45961dfbb8c9 · outbound

This paper cites Is space-time attention all you need for video understanding? In Proceedings of the International Conference on Machine Learning (ICML), 2021.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Is space-time attention all you need for video understanding? In Proceedings of the International Conference on Machine Learning (ICML), 2021

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:11.215819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.391202Z digest=sha256:9c21e4403e5c418ddc2d8f0a8955b484036424fc4c69c02fedcbd7c083ff0c2d

Observation a28ee2ca-3954-4bf7-8446-21d104855ea0 · outbound

This paper cites Executing your commands via motion diffusion in latent space.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Executing your commands via motion diffusion in latent space

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:11.202178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.395410Z digest=sha256:6230986279b0055e4ad2b7dbefda8a8f16ecaaf6617ff78aed6460928213d0b6

Observation 90ebcefb-4bc7-4367-b8d9-dee6fb004806 · outbound

This paper cites Synthesis of compositional animations from textual descriptions.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Synthesis of compositional animations from textual descriptions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:11.186824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.399775Z digest=sha256:bb8fd14c37c5a27b0892242616f49243194ac85b3194b31234de292605eb9e3f

Observation 49ff24c6-e9fa-438e-bc9f-b2d596e77a9b · outbound

This paper cites Generating diverse and natural 3d human motions from text.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Generating diverse and natural 3d human motions from text

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:11.171932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.403853Z digest=sha256:d8f09b2862a6004107fb0fce2d29c5f44ce2cccd2189ba8ec28eb8f3f8ff1f40

Observation 8a5d14a6-dfe0-438c-b474-2a6e28343603 · outbound

This paper cites Generating diverse and natural 3d human motions from text.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Generating diverse and natural 3d human motions from text

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:11.158129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.408426Z digest=sha256:b8d7b3feb27622ce2a7f4c2c8471fc56a62b90b00b417fb1fd46acbb58fa214a

Observation 333f7d69-1c8d-4870-84da-04d665e241a6 · outbound

This paper cites Tm2t: Stochastic and tokenized modeling for the reciprocal gener- ation of 3d human motions and texts.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Tm2t: Stochastic and tokenized modeling for the reciprocal gener- ation of 3d human motions and texts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:11.143537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.412816Z digest=sha256:57a293c66e80af987f7d8a8104936fa11ecf5a88346545a9c22c92d0af28c3ac

Observation fac05a42-127d-45b3-80aa-d334ecc73e0f · outbound

This paper cites Momask: Generative masked modeling of 3d human motions.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Momask: Generative masked modeling of 3d human motions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:11.127217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.417244Z digest=sha256:37b50b2e3a659063b710ddc041ee0db3d28cd3b6d1f1ab1321b37c433ef8cb51

Observation b369e979-c33e-4fa0-a6cf-66853cc81893 · outbound

This paper cites Como: Controllable motion generation through language guided pose code editing, 2024.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Como: Controllable motion generation through language guided pose code editing, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:11.112412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.422035Z digest=sha256:7e46284d4b1c87957a8c83c62c26d310798975aae03a9befc9536865837cef45

Observation e3d6f5a9-4afa-4b98-b874-45a64ad1f235 · outbound

This paper cites Motiongpt: Human motion as a foreign language.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Motiongpt: Human motion as a foreign language

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:11.096609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.426853Z digest=sha256:0ab9260161b26fed2418fa05ac81952836a005d386fd3d1639f07bc93a412ddc

Observation 2f6ec7da-fb1f-48a1-8e56-621faeb1751e · outbound

This paper cites Motionchain: Conversational motion controllers via multimodal prompts.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Motionchain: Conversational motion controllers via multimodal prompts

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:11.081415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.431443Z digest=sha256:3f370d83fc9f13751faed5c60231498a0c2fb02b346a14d47c6eda89ea969468

Observation c4e06e09-2702-4b59-98ba-a70402e57f1b · outbound

This paper cites Optimizing Diffusion Noise Can Serve As Universal Motion Priors.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Optimizing Diffusion Noise Can Serve As Universal Motion Priors

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:00:10.436091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:00:10.436091Z digest=sha256:9f448a06a4b01cc43dc3b892f753e163a604fa51ad3f02658b14336014824e6f

Observation f1c06346-6654-44ce-9dd7-9c2d71bec522 · outbound

This paper cites Guided motion diffusion for con- trollable human motion synthesis.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Guided motion diffusion for con- trollable human motion synthesis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:11.066277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.440767Z digest=sha256:1ff3771021e6ae27de95fb148d4841f22bf77a431785b0f1cfe1c2bbe442003c

Observation 1da931b3-a7ab-4d06-b902-71d2798a631e · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T19:00:10.444980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:00:10.444980Z digest=sha256:4ba9d8b3b497054a060c5030c3c7664cc6baaab577d5e88a9ea4319be8fad04f

Observation f6f8dac3-1ef8-4191-bb81-cd1b73a6eb75 · outbound

This paper cites Generating animated videos of human activities from natural language descriptions.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Generating animated videos of human activities from natural language descriptions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:11.042453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.449178Z digest=sha256:78730b823be1e62454e9eeba7b1096488a84e7cda06f563020a242a07a9e989a

Observation a03850cc-81e5-41b0-9eaa-38e7a74a8078 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Rouge: A package for automatic evaluation of summaries

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:00:10.453502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:00:10.453502Z digest=sha256:a225d3e1fb1c4d65377ab01fa53dd0d9ec8ef4b05dbb6654b8d111da1f74478a

Observation 141843eb-6f4f-4ee5-a6e8-9b539e19c3b9 · outbound

This paper cites Generation of Complex 3D Human Motion by Temporal and Spatial Composition of Diffusion Models.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Generation of Complex 3D Human Motion by Temporal and Spatial Composition of Diffusion Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:00:10.692286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.457731Z digest=sha256:cc8f33f2fdcdb5194678b1ea884ec3cff89e230df44b4dab208f232deb83280e

Observation 87b3f4a5-71d4-4044-8bad-15e66f26c1d9 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Bleu: a method for automatic evaluation of machine translation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T19:00:10.462696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:00:10.462696Z digest=sha256:50b7226c0ceccb9443e520b32b20bfe57c650f4d55b44f8b7f064e204642e9d3

Observation 955ab591-3b6f-49ba-9c9f-40cae5f4e8dd · outbound

This paper cites Temos: Generating diverse human motions from textual descriptions.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Temos: Generating diverse human motions from textual descriptions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:11.006847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.467177Z digest=sha256:f6d7487a38f407a803aa1cedee86b637083ffc9e38902e5cffef560d45858279

Observation 0be91046-a35e-442e-b02a-bc918f5739ad · outbound

This paper cites Mmm: Generative masked motion model.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Mmm: Generative masked motion model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.990856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.471566Z digest=sha256:5790d1deb435a9345bf4c35cff40a6f7115de3eea63bef143e950d47c73b3613

Observation 7331ff43-667a-44d7-91eb-75745c59761f · outbound

This paper cites Bamm: bidirectional autoregressive motion model.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Bamm: bidirectional autoregressive motion model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.975147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.475678Z digest=sha256:f0b0b1f8ca012957f88832a256d3ba01abbb23e0f317641f36754d4f337f00b8

Observation 2ebef19e-6f6b-4134-86f2-465e88160b9a · outbound

This paper cites Learning a bidirectional mapping between human whole-body motion and natural language using deep recurrent neural net- works.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Learning a bidirectional mapping between human whole-body motion and natural language using deep recurrent neural net- works

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.961099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.479906Z digest=sha256:56f7656534c3df4a4f5cb372bb23f6fe61eb3c2e987d5ffc96da93d4067929a9

Observation cd27aeae-af2b-471f-afd2-69677ccf29a5 · outbound

This paper cites Generat- ing diverse high-fidelity images with vq-vae-2.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Generat- ing diverse high-fidelity images with vq-vae-2

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.947450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.484061Z digest=sha256:aca9ceec3d3d2b39f39e53150f3d1efbe0fd27eef8d8a483f75884e2714a7b0e

Observation 03d927a9-4c31-4afc-b252-c7c2d4717aa6 · outbound

This paper cites Human motion diffusion as a generative prior.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Human motion diffusion as a generative prior

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.933975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.488765Z digest=sha256:1f8189225caaa09d5b3ac9a78e3e771ea224826d1648b2ce88f95498a5923305

Observation a8ae3c8b-7fed-48ea-ac4b-681386c292f8 · outbound

This paper cites MotionCLIP: Exposing Human Motion Generation to CLIP Space.

CoMA: Compositional Human Motion Generation with Multi-modal Agents MotionCLIP: Exposing Human Motion Generation to CLIP Space

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:00:10.493321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:00:10.493321Z digest=sha256:3a88df4591e1fc1a3a8fcd66ce1fe8214089e0264808bed18b19542c0bdb404c

Observation a9978507-2cc8-4a14-a4c4-e7bb11544156 · outbound

This paper cites Human motion diffusion model.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Human motion diffusion model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.920196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.498472Z digest=sha256:c79cf7a47a225e076ce38e8296069047912aeea9eefaefa93a89b95e715d703f

Observation 6be87e95-b4ba-42c0-b601-d452ede9848a · outbound

This paper cites Nvae: A deep hierarchical variational autoencoder.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Nvae: A deep hierarchical variational autoencoder

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.906681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.503617Z digest=sha256:5c260cb885bc0bb5e13d17e771356af234e03c1cce4f592f84de41c6ce3b0c40

Observation 2ed343e7-210d-4431-ae9e-eb3e99d7cba9 · outbound

This paper cites Neural discrete representation learning.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Neural discrete representation learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.892236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.508568Z digest=sha256:b0f972afd3989afb1ffa76ade8fa0de0cff580feb40eeeb0647ad7bf00e6b9bb

Observation afd2a95d-a97a-4d31-8e10-572485c0f37a · outbound

This paper cites Neural discrete representation learning.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Neural discrete representation learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.877641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.512972Z digest=sha256:f24b3c3f8163a952c0eaf13e7ff7ae1ba3eab4708dd06788af7df9f328b896fd

Observation ac324ba9-a799-4086-bdbc-b85c0fae5c2f · outbound

This paper cites Cider: Consensus-based image description evalu- ation.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Cider: Consensus-based image description evalu- ation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.863003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.517395Z digest=sha256:ca3124fd287dbf6af86023e348766512a01b3b59396bf717cb6884fcd3a954fc

Observation 239b5771-0720-455a-b6c0-132e85953186 · outbound

This paper cites Fg-t2m: Fine-grained text-driven human motion generation via diffusion model.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Fg-t2m: Fine-grained text-driven human motion generation via diffusion model

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.847907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.522117Z digest=sha256:6e65adba7db61efca9e51e5dc53c530f2c65a3477f191a13194a85f638baf6c5

Observation 27900b70-bcb0-4693-8446-9af75264d895 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

CoMA: Compositional Human Motion Generation with Multi-modal Agents InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T19:00:10.526296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:00:10.526296Z digest=sha256:02f47aa11fef0ab3c166ec443143c39bc8b91a1d4eed3502584b1fea6c22e4ed

Observation 58033f15-761b-4826-b236-ee34b50c189d · outbound

This paper cites Chain-of- thought prompting elicits reasoning in large language models.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Chain-of- thought prompting elicits reasoning in large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.831485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.530766Z digest=sha256:0182e638e2935ca5aa6476643c16537289ca65ea31b226fe6c8e4a9bc8948269

Observation 8f2c2547-68c1-4278-b4ce-95f4de278728 · outbound

This paper cites Motion-agent: A conversational framework for human motion generation with llms, 2024.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Motion-agent: A conversational framework for human motion generation with llms, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.815781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.535283Z digest=sha256:4dc9fdd63e61c5386ebd904fe11ab0f3f371f3db9fe15d6b074bd1189ff5711e

Observation dacbca06-3457-4f17-8320-4535ba01713f · outbound

This paper cites Omnicontrol: Control any joint at any time for human motion generation.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Omnicontrol: Control any joint at any time for human motion generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.798718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.539479Z digest=sha256:239cc43baccfccc5bb6ece8c30d49c67739715e514bc9b2ecaf9705a7a555308

Observation 03658c26-6736-4b8a-b14d-4e197e7c89e5 · outbound

This paper cites Attribute2image: Conditional image generation from visual attributes.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Attribute2image: Conditional image generation from visual attributes

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.784885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.543557Z digest=sha256:66a405566826e2003d89be67fcbf278ec28be606ed0e997c5c63ca995bd93c81

Observation d527f119-4457-4763-8c83-52fa0b28fced · outbound

This paper cites Soundstream: An end-to- end neural audio codec.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Soundstream: An end-to- end neural audio codec

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.770124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.548141Z digest=sha256:c9f3b498cc98d18c12f0fd6483f1069d42c42c60ace7932df58b22461f4c02f4

Observation 88aa3ae7-21fc-48a9-938d-84332972f880 · outbound

This paper cites T2m-gpt: Generating human motion from textual descriptions with discrete representations.

CoMA: Compositional Human Motion Generation with Multi-modal Agents T2m-gpt: Generating human motion from textual descriptions with discrete representations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.755555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.552114Z digest=sha256:b469ca1459012138bd315e783861c09b067c62797287306eb3beef7715a974ac

Observation 8b2cf08d-c217-4932-98ed-0c751abd0167 · outbound

This paper cites MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model.

CoMA: Compositional Human Motion Generation with Multi-modal Agents MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T19:00:10.556576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:00:10.556576Z digest=sha256:e71de568f81158e5b08b552757e2e6b561a6175401eb9ad2e0c117d5f378598c

Observation 85e193bd-00fb-4747-8761-a01976b33fb2 · outbound

This paper cites ReMoDiffuse: Retrieval-Augmented Motion Diffusion Model.

CoMA: Compositional Human Motion Generation with Multi-modal Agents ReMoDiffuse: Retrieval-Augmented Motion Diffusion Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T19:00:10.561526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:00:10.561526Z digest=sha256:aa04d9e95f46771ba8411bd8d2751eb8ea255b0e61d5aafcd1ada48e3b3bc2b0

Observation 6b02d67c-90d1-485f-bd03-ee19956c6023 · outbound

This paper cites Finemogen: Fine-grained spatio- temporal motion generation and editing.

CoMA: Compositional Human Motion Generation with Multi-modal Agents Finemogen: Fine-grained spatio- temporal motion generation and editing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:00:10.739514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:00:10.566685Z digest=sha256:32f25db8878d9dc79bbd9f8f16633b7f81acc7a7d291f4f6fa1bbbcb837dfb56

Observation 7f6a5162-2dac-4a3e-8aaf-e61805966484 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

CoMA: Compositional Human Motion Generation with Multi-modal Agents BERTScore: Evaluating Text Generation with BERT

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T19:00:10.571415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:00:10.571415Z digest=sha256:efcd1859254101250696c504ca9a9a9f0804bd1aaf5a980dd240b18605a2cbe3

Pith citing papers

Observation 38885931-82a8-49b5-9d59-d0ff84385ed7 · inbound

Absolute Coordinates Make Motion Generation Easy cites this paper.

Absolute Coordinates Make Motion Generation Easy CoMA: Compositional Human Motion Generation with Multi-modal Agents

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:20.685328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:20.685328Z digest=sha256:cbe0866bba4beaea0c46e2e5bc01dbb129efc3380a8c420458c11619c9d30771

Observation 2d162210-87f4-4338-b20a-c50fb13eafb4 · inbound

Multi-Modal Manipulation via Multi-Modal Policy Consensus cites this paper.

Multi-Modal Manipulation via Multi-Modal Policy Consensus CoMA: Compositional Human Motion Generation with Multi-modal Agents

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:51:20.246560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T11:50:20.766891Z digest=sha256:c5512b877c0b75b8f009ce24ee1d6dd68ecf15def05e17bee1aa4cf138df55af

Observation 728fbaf6-d3bd-44f2-a166-9c0415fd05ea · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning CoMA: Compositional Human Motion Generation with Multi-modal Agents

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:32.650256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:32.650256Z digest=sha256:4a0c5225835085f3066e55fd2da7c7b673eeaf91d6e793f6023829c9da66a3f7

Observation 7bd5556f-4023-411f-9219-28f0c3fa9a5c · inbound

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents cites this paper.

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents CoMA: Compositional Human Motion Generation with Multi-modal Agents

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:58:31.799304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T20:56:58.770875Z digest=sha256:2680e581c47e2ea513e8b74911ddcbd28b5522610ccd2563081d557c01f44169

Observation c6481434-0f9e-4afa-9cc2-cdf53114e0fb · inbound

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision cites this paper.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision CoMA: Compositional Human Motion Generation with Multi-modal Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.588885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.588885Z digest=sha256:ab1c61146f7173ebc67010245e0c73b1568abea33dbf59545f35ccdda7c809b0