{"id":"bbd783f0-3262-4f5f-8b2a-d0cc6be9f4ab","arxiv_id":"2506.03191","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey paper reviews multimodal generative AI and autoregressive LLMs for text-driven human motion generation, with comparative tables of models, datasets, and metrics.","lead":"This preprint is a survey of text-to-motion AI models, reviewing how large language models, diffusion models, GANs, and VAEs generate human motion from text descriptions. It compiles architectures, datasets, evaluation metrics, and applications, and claims to be the first review to cover human motion understanding and generation together.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 6's quantitative comparison lists Precision values exceeding 1 (e.g., AlertMotion 5.400, ActFormer 2.550, UDE 8.210), contradicting the paper's own Table 5 bounds (0≤Precision≤1); the survey's central benchmark table is not reliable as printed.","rationale":"The reader's weakest assumption is that Table 6's comparative results are accurate transcriptions and comparable across models. My independent check found direct evidence supporting that assumption failure: Precision values greater than 1 contradict the paper's own Table 5 bounds, and at least one row is citation-mismatched. This is the most load-bearing concern because the survey's distinct value proposition is comprehensiveness plus reliable cross-method comparison; if the centerpiece comparison table is corrupted, the survey cannot function as a dependable reference. The errors are concrete and localizable, so the appropriate disposition remains CONDITIONAL — the paper should be accepted only after the table is corrected and re-verified, not rejected outright. I therefore keep the reader's verdict unchanged. The duplicate Section 6/7 text and other typos are secondary; they affect polish but not the central argument. My recommendation aligns with the reader, and the concrete test would settle whether the impossible values are transcription errors or a hidden metric redefinition.","tokens_in":44075,"tokens_out":2026,"duration_ms":22973,"concrete_test":"Locate the HumanML3D evaluation numbers in the original papers cited for AlertMotion [82], ActFormer [92], and UDE [173]. If their Precision entries (5.400, 2.550, 8.210) appear verbatim in the sources, determine the metric definition used (e.g., percentage or a differently normalized R-Precision) and state it alongside Table 6; if they do not appear, correct the entries or remove them. Additionally, check whether reference [103] (Ghosh et al., 2021, GAN-based compositional animation) is genuinely the source of the row labeled 'VQ-VAE Mot'; if not, re-cite the correct VQ-VAE paper. The table passes only if every row matches its cited source after normalization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is to be the first comprehensive, unique survey covering all of HMUG, and its added value over prior surveys depends on trustworthy cross-method comparison. That comparison is carried by Table 6, yet Table 6 reports Precision values of 5.400 (AlertMotion), 2.550 (ActFormer), and 8.210 (UDE) on HumanML3D, while Table 5 in the same paper defines Precision as bounded by 0 and 1 with no percentage caveat; the MM column is also left blank for several rows. There are also citation mismatches: the row labeled 'VQ-VAE Mot [103]' points to reference [103], Ghosh et al., 'Synthesis of Compositional Animations from Textual Descriptions,' which is a GAN-based text-to-motion work, not a VQ-VAE method, and 'Unify MoGPT [146]' appears to conflate separate unified-model papers. Because Table 6 is the only quantitative evidence for the survey's comparative value, any transcription error or metric mismatch propagates to every downstream statement about which architecture family performs best. The internal contradiction with Table 5 means the numbers cannot be correct as printed, so the central benchmark is unreliable until re-verified against the original sources.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of text-conditioned human motion understanding and generation using multimodal generative AI and autoregressive large language models. It reviews data representations, text and motion tokenizers, LLM-based motion models, GAN/VAE/diffusion text-to-motion methods, unified frameworks, datasets, evaluation metrics, applications, and open challenges. The authors claim that this is the first survey to cover all areas of Human Motion Understanding and Generation (HMUG), and they support the survey with comparative tables, including a quantitative comparison of methods on HumanML3D and KIT-ML.","tokens_in":44321,"tokens_out":6289,"duration_ms":57238,"significance":"If the comparative tables and citation mappings were accurate, this survey would be a useful reference because it gathers a large body of recent work and organizes it by architectural family and task. The breadth is real: the paper reviews more than 250 references and provides useful taxonomy tables and figures that map models, backbones, datasets, and tasks. However, the survey's central value as a cross-method benchmark depends on Table 6 and on the traceability of Tables 2-4, and those components are currently not reliable as printed. The paper is a survey, so it contains no machine-checked proofs or code, but its factual claims about other papers must be verifiable against the cited sources; at present they are not.","major_comments":[{"comment":"Table 5 defines Precision with bounds 0≤Precision≤1 and MM Score with bounds 0≤MMS≤1, but Table 6 reports Precision values of 5.400 (AlertMotion), 2.550 (ActFormer), and 8.210 (UDE) on HumanML3D, and MM values above 1 in nearly every row. Since Table 6 is the only quantitative cross-method comparison in the survey, these internally inconsistent numbers cannot support the paper's comparative conclusions. The authors must re-extract every value from the original papers, state the metric definition used, and ensure consistency with Table 5.","section":"Table 5 and Table 6"},{"comment":"Several rows in Table 6 do not point to the cited work. The row 'VQ-VAE Mot [103]' cites [103], Ghosh et al., 'Synthesis of Compositional Animations from Textual Descriptions,' which is not a VQ-VAE method; 'Unify MoGPT [146]' cites [146], Ma et al., which is the MoFusion diffusion paper; and 'WalkLLM [221]' cites [221], PackDiT, not the WalkLLM pedestrian-motion paper. In addition, there are two distinct papers named MotionLLM ([62] and [74]), and Table 6 lists MotionLLM [62] twice with different FID and Precision values. The reference-to-row mapping must be corrected throughout the tables.","section":"Table 6 and reference list"},{"comment":"Sections 6 and 7 have identical titles ('Datasets and Evaluation Metrics') and identical opening paragraphs, with Section 7 containing the actual subsections while Section 6 is left empty in substance. This duplication is a structural error that must be fixed by merging the content into a single section and renumbering the subsequent sections before the survey can be read as a coherent document.","section":"Sections 6 and 7"},{"comment":"The claim that this is 'the first, and unique attempt that covers all the areas of Human Motion Understanding and Generation (HMUG)' is not substantiated by Table 1 as presented. The checkmark criteria in Table 1 are undefined, and several prior surveys (e.g., [22] and [31]) overlap substantially with the stated scope. The authors should either provide an operational definition of 'all areas of HMUG' and demonstrate specific coverage gaps relative to each prior survey, or soften the uniqueness claim.","section":"Introduction and Table 1"}],"minor_comments":[{"comment":"Equation (14) is written as FID(P1,P2)^2 = ..., while the surrounding text refers to 'the FID score'; please adopt the standard convention FID = ||μ1-μ2||^2 + Tr(Σ1+Σ2-2(Σ1Σ2)^{1/2}) to avoid ambiguity.","section":"Equation (14)"},{"comment":"Table 3 lists 'ActFormer [95]' and 'FineMoGen [93]', but the cited references are the DCGAN paper and 'To Create What You Tell', respectively; the correct citations appear to be [92] and [89]. Please verify and correct these mappings.","section":"Table 3"},{"comment":"The abstract and Section 1 state that the survey focuses exclusively on text and motion modalities, but Section 7.3 and Table 4 include audio, speech, music, and emotion datasets and methods; please reconcile the stated scope with the included material.","section":"Scope, Abstract and Section 7.3"},{"comment":"Table 2 uses the overlapping names 'MotionGPT-3 [78]', 'MotionGPT-2 [137]', and 'T2M-GPT [77]' while the text and reference list use similar names for different papers; please adopt a consistent naming and citation scheme so that readers can map each row to its source.","section":"Table 2"},{"comment":"The FID and other metric values in Table 6 are reported without the evaluation protocol used (e.g., number of samples, text prompt set, and whether metrics come from the original papers or from re-evaluation), so even after correcting the transcription errors the values may not be directly comparable across methods; please state the protocol explicitly.","section":"Table 6 and evaluation protocols"}],"recommendation":"major_revision","confidential_remarks":"The manuscript needs a systematic audit of all tables and reference mappings, not just spot fixes. The duplicate Sections 6 and 7 and the impossible metric values in Table 6 suggest the paper was assembled too quickly. If the authors correct these issues and re-verify Table 6 against original sources, the survey could become a useful contribution; if the quantitative table cannot be repaired, I would not support publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know two things about this paper. It is a broad survey of text-to-motion generation that really does try to cover understanding and generation together, with LLM and diffusion approaches, and a lot of recent work is packed into useful tables. But the current version has internal errors that make it unreliable as printed, and the worst are in the central quantitative comparison.\n\nThe survey's value is organizational. It maps the text-to-motion landscape, breaks down data representations, tokenizers, generative backbones (GANs, VAEs, diffusion, autoregressive LLMs), and unified models. The tables that list methods with their backbones, datasets, and tasks are genuinely handy. The equations for VQ-VAE, VQ-GAN, diffusion, and the evaluation metrics are stated formally enough to be useful. For someone entering the subfield, this is a reasonable place to start.\n\nThe soft spots are not minor. Sections 6 and 7 are duplicated verbatim for the opening paragraphs, so the paper has not been through any careful edit. Table 6, the only cross-method benchmark comparison, lists Precision values of 5.400, 2.550, and 8.210 on HumanML3D, while Table 5 in the same paper defines Precision with bounds 0≤Precision≤1. Those numbers cannot be correct as printed. The same table has citation mismatches: the row 'VQ-VAE Mot [103]' points to a GAN-based animation paper. These are not one-off typos; they undermine the paper's stated added value as the definitive cross-method comparison. There is also an internal inconsistency in the MMS formula (Eq. 18) which need not be bounded by 1, yet Table 5 bounds it between 0 and 1.\n\nI would not cite Table 6 or the MMS bound without checking every original source. If the authors re-verify all transcriptions and fix the duplication, the survey could be a solid reference. In its current form, I would not trust the numbers.\n\nThis paper is for researchers who want a structured overview of text-to-motion methods and a starting bibliography. For that purpose, the breadth is useful. It deserves a serious referee, but the revisions should be mandatory: fix the duplicated sections, re-verify every value in Table 6 against the original papers, correct the citation labels, and soften the claimed uniqueness.\n\nRecommended action: send to peer review, but expect major revision.","headline":"A useful survey of text-to-motion with a broad scope, but its central comparison table contains impossible numbers and the paper is not trustworthy as printed.","tokens_in":44813,"tokens_out":3148,"would_cite":false,"duration_ms":30135,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims to be the first to map all of human motion understanding and generation, organizing text-conditioned synthesis into autoregressive LLMs, diffusion models, GANs, VAEs, and unified AR-diffusion frameworks.","keywords":["human motion generation","text-to-motion","multimodal large language models","diffusion models","autoregressive models","motion understanding","generative AI survey","unified motion framework"],"falsifier":"A reader can settle the uniqueness claim by checking the surveys listed in the paper's own Table 1: if any of them already covers all nine scope columns the paper counts for itself, the 'first and unique' assertion is false. The benchmark tables are independently falsifiable, since recomputing entries against the cited papers will expose an error whenever a reported Precision exceeds the [0,1] bounds the paper itself lists (for example, the 5.400 value in Table 6).","tokens_in":43887,"feed_emoji":"🏃","tokens_out":7146,"duration_ms":66674,"temperature":0.7,"pith_summary":"This paper argues that human motion understanding and generation (HMUG) is best viewed as a multimodal problem in which text conditions motion, with two main engines: autoregressive large language models and diffusion models. It claims to be the first survey to cover all of HMUG in one place—data representation, generation architectures, unified AR+diffusion frameworks, datasets, evaluation metrics, and applications—where earlier surveys covered only subtopics such as prediction, image or video generation, or generative methods without multimodal LLMs. The practical value the authors see is that a researcher entering text-to-motion can use this survey to choose an architecture, dataset, and metric without assembling the literature from scattered papers. The paper also argues that the next step is a unified model that plans, understands, and generates motion, together with a unified evaluation metric to compare such models fairly.","feed_headline":"One survey maps text-to-motion AI from LLMs to diffusion","feed_subtitle":"Covers architectures, datasets, metrics, and applications for making machines move from text prompts.","key_machinery":"The organizing device is an architecture taxonomy: autoregressive LLMs, which predict the next token and therefore need motion converted into discrete tokens by codebook-based vector quantizers such as VQ-VAE and VQ-GAN; diffusion models, which add and remove Gaussian noise in a forward-reverse process that ideally runs in a compressed latent space; and unified frameworks that align text and motion in a shared embedding space, using a multimodal transformer or a mixture-of-experts connector to let the autoregressive and diffusion paradigms cooperate. Each mechanism does a specific job in the survey: it explains why methods behave differently, since autoregressive models capture long-range text-motion dependencies but lose fine detail at tokenization, diffusion models produce high-fidelity motion at the cost of many denoising steps, and unified models aim to get both understanding and generation from a single training objective.","core_discovery":"In the authors' telling, the central discovery is that text-conditioned human motion generation has converged on two complementary paradigms—autoregressive LLMs that treat motion as a foreign language of discrete tokens (via VQ-VAE or VQ-GAN tokenizers) and diffusion models that refine noisy motion in continuous or latent space—and that recent work is combining them into unified frameworks. The survey's own contribution is the claim of uniqueness: no prior survey, it says, covers both motion understanding and generation together with multimodal generative AI and autoregressive LLMs, because previous reviews stop at prediction, at image or video generation, or at general generative models without LLMs. It supports this claim by categorizing methods by architecture and backbone, comparing them on the HumanML3D and KIT-ML benchmarks with fidelity, diversity, and consistency metrics, and mapping them onto applications from healthcare to autonomous driving.","pith_inferences":["If discrete motion tokens become standard, general-purpose LLM tooling—instruction tuning, prompting, retrieval, and even adversarial red-teaming—would apply to motion almost directly, which is why the survey's unified AR+diffusion direction is a plausible successor to diffusion-only text-to-motion models.","A concrete testable extension of the survey's call for a unified metric would be to build a single score that normalizes fidelity, diversity, and consistency across HumanML3D and KIT-ML and validates it against human perceptual judgments.","The taxonomy implies that motion understanding and generation are converging: once the same tokenizer feeds both an LLM and a diffusion decoder, captioning and synthesis become two directions of one mapping, an idea the survey hints at but does not develop.","The paper's comparison tables are best treated as a starting point to be checked against the original sources before being reused as a benchmark, since cross-method comparability is assumed rather than demonstrated."],"forward_implications":["A newcomer to text-to-motion generation can select an architecture—autoregressive LLM, diffusion, GAN, VAE, or unified—directly from the survey's taxonomy, with matching datasets and metrics.","Autoregressive LLM approaches that treat motion as discrete tokens are mature enough to generate, caption, retrieve, and reason about motion, not merely synthesize sequences.","Unified AR+diffusion models trained in a shared embedding space with a multimodal transformer or mixture-of-experts connector are presented as the direction that combines understanding and generation.","The absence of a standardized evaluation framework is identified as a real bottleneck, and the survey argues for a unified metric that combines fidelity, diversity, and consistency, citing work toward that goal.","Real-world deployment in rehabilitation, VR/AR, gaming, robotics, surveillance, and autonomous vehicles depends on making motion generation efficient, controllable, and long-form, which the survey lists as the main open challenges."],"supporting_citations":[{"why":"The prior human-motion-generation survey the paper positions itself against, because it excludes multimodal LLMs and transformers.","marker":"[21]"},{"why":"The prior multimodal generative-AI survey it extends, which covers text, audio, video, and images but does not cater specifically to human motion generation.","marker":"[22]"},{"why":"MotionDiffuse, the diffusion-based text-to-motion model that anchors the survey's diffusion section.","marker":"[5]"},{"why":"T2M-GPT, the discrete tokenization method that lets motion be treated as a foreign language for autoregressive LLMs.","marker":"[80]"},{"why":"MotionLLM, the multimodal motion-language model the survey uses to ground its discussion of understanding, captioning, and generation with LLMs.","marker":"[62]"},{"why":"AvatarGPT, the all-in-one motion understanding-planning-generation framework cited as a milestone toward unified models.","marker":"[76]"},{"why":"The latent-space motion diffusion model that anchors the survey's efficiency comparisons and its discussion of latent diffusion.","marker":"[10]"},{"why":"The prior work on a unified evaluation framework that the survey relies on for its claim that motion metrics need standardization.","marker":"[32]"}],"fun_headline_variants":["Survey finds text-to-motion merging LLMs and diffusion","Text-to-motion AI converges on LLMs and diffusion","Survey unites LLMs and diffusion for text-to-motion","How LLMs and diffusion shape text-driven motion","Two paradigms converge: LLMs and diffusion for motion"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the comparison-table numbers copied from the cited papers are accurate and were measured under comparable protocols; if they are wrong or incomparable, the survey's value as a cross-method benchmark is compromised.","fun_headline_variants_meta":{"raw":{"variants":["Survey finds text-to-motion merging LLMs and diffusion","Text-to-motion AI converges on LLMs and diffusion","Survey unites LLMs and diffusion for text-to-motion","How LLMs and diffusion shape text-driven motion","Two paradigms converge: LLMs and diffusion for motion"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001243,"raw_usage":{"total_tokens":5094,"prompt_tokens":933,"completion_tokens":4161,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":4082}},"tokens_in":549,"tokens_out":4161,"duration_ms":28850,"temperature":1.0,"reasoning_tokens":4082,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:02:44.882467+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader can settle the uniqueness claim by checking the surveys listed in the paper's own Table 1: if any of them already covers all nine scope columns the paper counts for itself, the 'first and unique' assertion is false. The benchmark tables are independently falsifiable, since recomputing entries against the cited papers will expose an error whenever a reported Precision exceeds the [0,1] bounds the paper itself lists (for example, the 5.400 value in Table 6).","supporting_citations":[{"cited_title":"Multi-Modal Generative AI: Multi-modal LLM, Diffusion and Beyond,","cited_arxiv_id":null,"evidence_quote":"The prior multimodal generative-AI survey it extends, which covers text, audio, video, and images but does not cater specifically to human motion generation."}],"review_version":1}