{"id":"aa11058e-7628-449f-92c6-97168087f7d2","arxiv_id":"2608.06722","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"CustomDance combines an MLLM-based choreographic planner, multimodal dance-phrase retrieval, and diffusion inpainting into one three-stage interactive system for user-customized 3D dance generation.","lead":"This paper presents CustomDance, an interactive system that plans, composes, and refines 3D dance choreography from music and user text preferences through three stages. It matters because it directly targets the gap between user creative intent and automatic music-to-dance generation with an authoring workflow.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2's objective-metric advantage is confounded by retrieval from the real-motion dance library and by human-guided phrase selection, so the quantitative superiority claim is not isolated to the proposed system.","rationale":"I read the paper as a systems-and-evaluation contribution rather than a derivation, and I agree with the reader's CONDITIONAL verdict: the user-authoring studies are internally coherent, the ablations show plausible benefits for each stage, and the three-stage workflow is genuinely novel for HCI. The most load-bearing weakness is not any internal inconsistency in the pipeline equations; it is the interpretability of the objective quality comparison. CustomDance is a retrieval-plus-inpainting system whose retrieved phrases come from a curated library of real motion capture, so its output distribution is anchored to real dance motion in a way that automatic generators are not. The FID/Diversity/BAS numbers in Table 2 consequently blend three effects: the library's motion distribution, the user's ability to choose high-quality candidates from a top-10 list, and the actual generative inpainting/refinement. Only the last of these is attributable to the proposed architecture, and the paper does not provide an ablation that separates it. The limitations paragraph concedes exactly this concern, which strengthens rather than weakens the case that it is load-bearing. The reader's weakest_assumption focused on the cleanliness of the dance library; this is related but distinct, so I mark partial agreement. A retrieval-only baseline with error bars is a concrete, feasible check that would settle whether the objective superiority claim survives. If it does not survive, the user-study evidence still supports the authoring-workflow claims, so the verdict remains CONDITIONAL rather than REJECT: the authors should either report the retrieval-only ablation and cross-dataset FID, or substantially soften the objective-comparison claim in the abstract and introduction.","tokens_in":17491,"tokens_out":5836,"duration_ms":62015,"concrete_test":"Recompute Table 2 for a Retrieval-Only baseline: use the same slot anchors and the same curated library, select the top-1 retrieved phrase per slot without MLLM cues, Diagnoser, or diffusion inpainting, and concatenate the phrases with simple linear blending. If Retrieval-Only yields FID_k and BAS statistically indistinguishable from CustomDance's 22.55 and 0.233 (e.g., overlapping 95% confidence intervals across at least 10 independent runs or participant selections), then the objective-metric advantage is attributable to the library and retrieval, not to the proposed coarse-to-fine workflow.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim — that CustomDance achieves better dance-quality metrics than Lodge, MEGADance, and a baseline editor (Table 2: FID_k 22.55, Div_k 7.02, BAS 0.233) — is not a clean test of the system. CustomDance outputs are mostly user-selected 4-second phrases taken from a curated library of real motion capture (7.7h FineDance plus ~3h HMR-reconstructed video, §7.1), with diffusion inpainting filling gaps and repairing localized artifacts. FID, Diversity, and BAS are therefore computed on fragments that lie inside, or very close to, the real-motion distribution used for evaluation, while Lodge and MEGADance must synthesize those fragments from noise or learned priors. Section 7.3 states that a repair workflow is applied to Baseline, Lodge, and MEGADance before evaluation, but that workflow cannot supply the same library retrieval or user-guided selection. The paper's own Limitations section (§8) concedes that \"fidelity-oriented metrics may partially favor retrieval-supported outputs.\" Because the headline claim explicitly includes the FID 22.55 / BAS 0.233 numbers, this retrieval-and-selection confound is load-bearing: it means Table 2 does not establish that the MLLM planning, multimodal Retriever, or diffusion Completer/Remaker improve objective fidelity. The authoring-time and perceived-quality results in Study 1 are less affected by this concern, since they compare interface conditions within the same retrieval library, but the cross-method objective comparison remains unvalidated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"CustomDance presents a three-stage interactive choreography system: an MLLM (Gemini) generates temporal anchors and creative cues from music and a user text prompt; a trimodal retriever ranks 4-second phrases from a 10.7-hour dance library conditioned on local music, text, and joint-group intensity/variety controls; and a diffusion inpainter fills gaps and repairs localized artifacts. The paper reports a 25-participant within-subject authoring study (Study 1), a blinded 25-rater quality ranking against Lodge, MEGADance, and a baseline editor (Study 2), and a 90-second long-sequence study with expert raters and professional scratch choreography (Study 3), along with component-level evaluations. The central claim is that the full system improves authoring efficiency, perceived control, and final dance quality over the ablated conditions and over end-to-end generation baselines.","tokens_in":17744,"tokens_out":5239,"duration_ms":49013,"significance":"If the results hold, the paper is a meaningful contribution to AI-assisted choreography and interactive character animation: it provides a concrete instantiation of the choreographer-inspired workflow, with a clear interface design and a thorough user-study protocol. Strengths include the within-subject ablation design, the blinded ranking procedure, the long-sequence comparison that includes professional scratch choreography, and the honest Limitations section that already concedes the retrieval-based metric concern. The authoring-time reductions (11.66 vs 28.41 minutes) and high rank-1 rates are impressive and support the central authoring-benefit claim. However, the objective metric comparison and the artifact-count measure have confounds that must be resolved before the quantitative superiority claims can be accepted as stated. The circularity burden of the paper is low because it makes no formal derivation, and the free parameters (slot duration, candidate list size, CFG scale, dropout probabilities) are disclosed explicitly.","major_comments":[{"comment":"The objective comparison is confounded by retrieval and human selection. CustomDance outputs are composed largely of user-selected motion-capture phrases from a curated library, while Lodge and MEGADance synthesize from learned priors. FID, Diversity, and BAS therefore measure different objects: retrieval-aided composition versus full generation. The paper's own Limitations section (§8) concedes that 'fidelity-oriented metrics may partially favor retrieval-supported outputs.' Since the abstract and Section 7.3.2 explicitly claim superiority based on these numbers, Table 2 does not establish that the proposed generative and planning components improve motion fidelity. I recommend either reporting metrics on the diffusion-inpainted (gap and repair) portions only, or giving Lodge and MEGADance access to the same retrieval library, or re-labeling the comparison as a system-level study rather than a method-level one.","section":"§7.3, Table 2"},{"comment":"The artifact-count metric appears circular with the Diagnoser. The artifact count is defined as 'thresholded kinematic anomalies,' and the Diagnoser uses the same class of AKE/RKE kinematic proxies. Unsurprisingly, conditions with the Diagnoser enabled report 0.00 artifacts. The footnote that the offline metric is 'separate from the AKE/RKE curves' is not sufficient, because both target the same kinematic outliers (root jumps, limb twists). Please provide an independent artifact definition (e.g., foot skating, penetration, joint-angle limit violations) or demonstrate that the offline metric was not guiding the repair actions in a way that trivially removes exactly those anomalies.","section":"§7.2.1–§7.2.2, Table 1"},{"comment":"No error bars, confidence intervals, or significance tests are reported for any objective metric. It is unclear how many generations or seeds were used and whether the FID, Diversity, and BAS values are statistically distinguishable. Please report mean and standard deviation over at least five independent runs (or over the five selected dances) and apply a statistical test when comparing methods.","section":"§7.3, Table 2"},{"comment":"The dance library combines 7.7 hours of FineDance motion capture, approximately 3 hours of PromptHMR-reconstructed video, and StableMotion augmentations, with no ground-truth validation of the reconstructed or augmented motions. If HMR reconstruction noise, coordinate misalignment after the FineDance alignment, or augmentation artifacts enter the library, both the retrieval quality and the objective metrics in Table 2 would be affected independently of the interaction design. The paper should report validation (e.g., reconstruction error on a held-out set, or manual inspection rates) or an ablation showing that retrieval quality and objective metrics are stable to the HMR/augmented subset.","section":"§7.1"}],"minor_comments":[{"comment":"Equation (5) introduces the masked denoising update, but the notation q(x_known, t-1) is not defined; please clarify the forward-process conditional and the noise schedule used.","section":"§6.3, Eq. (5)"},{"comment":"Table 3 reports FSR and Jitter without in-text definitions; a brief definition or a precise pointer to the original sources would help the reader interpret the long-sequence results.","section":"§7.4, Table 3"},{"comment":"The artifact-count metric is only described as 'thresholded kinematic anomalies' with details deferred to the supplementary material; a concise definition in the main text is needed to assess the circularity concern raised in the major comments.","section":"§7.2.1"},{"comment":"The 'same repair workflow' applied to Baseline, Lodge, and MEGADance is not specified; please describe the repair protocol, including who invoked it and what constraints were observed, so that the fairness of the comparison can be evaluated.","section":"§7.3.1"},{"comment":"The behavior of overlapping slots ('Selecting one temporarily invalidates conflicting slots') is confusing; please clarify whether the conflicting slot's content is preserved and how the user's undo action restores it.","section":"§4"},{"comment":"Some references are not yet published (e.g., MEGADance [NeurIPS'25], OmniDance [2026b]); if final versions exist, they should be cited to help readers verify the baselines.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper's technical novelty is moderate since it combines known components (MLLM planning, contrastive retrieval, diffusion inpainting), but the integration and evaluation are substantial. The main concern is that the objective metric claims are not yet clean: the retrieval confound in Table 2 and the apparent circularity of the artifact-count metric need to be resolved before the quantitative superiority claims can be accepted. The reliance on a closed API (Gemini) and on supplementary material for key definitions (retriever training, artifact metric, repair protocol) also complicates independent verification. The paper would be considerably stronger if the retrieval confound were addressed head-on and if the artifact metric were independently defined."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this one. First, the system is real and the interaction design is the contribution: MLLM anchor/cue planning, multimodal retrieval of 4-second phrases, and diffusion inpainting for gaps and localized repair. That three-stage split is genuinely new as a workflow, and the ablation study gives it real support. Second, the headline objective metrics in Table 2 are not a clean comparison, and the paper's own Limitations section more or less admits it. That doesn't sink the paper, but it should stop you from citing the FID 22.55 / BAS 0.233 numbers as evidence that the system is a better generator.\n\nThe user studies are the strong part. Study 1 is a sensible within-subject ablation with 25 participants, counterbalanced, with Holm-corrected nonparametric tests. The authoring-time reduction from 28 to 11 minutes and clip-replacement drop from 20 to 5 are believable effects of retrieval ranking. The Diagnoser's artifact counts being zero in Diagnoser conditions is plausible, though the metric is partly circular since the offline artifact counter flags the same kinematic anomalies the Diagnoser visualizes. Study 3 is ambitious: 90-second comparisons against automated generators, non-professional scratch, and professional scratch choreography captured with mocopi. CustomDance lands near the professional scratch on FSR and Jitter and beats the automated baselines on expert ratings. That is a meaningful result for an interactive authoring tool.\n\nThe soft spots beyond the retrieval confound: no error bars or significance tests on Table 2; the comparison mixes interactive and non-interactive regimes even after the shared repair workflow; and no code or data released. The 10.7-hour library includes ~3 hours of HMR-reconstructed video, so library contamination is a real concern for both retrieval quality and the objective metrics. These are fixable with released artifacts and a clarified protocol. The closed Gemini dependence is a practical limitation, not a scientific one.\n\nWho benefits? HCI and graphics researchers working on human-in-the-loop motion authoring, and anyone building on the choreographer-workflow analogy. It deserves a serious referee. I'd send it out, but I'd ask the authors to either present Table 2 as a systems comparison rather than a method-level head-to-head, or to add an ablation that isolates the generative contribution from the retrieval advantage.","headline":"A well-built interactive choreography system with a genuine three-stage workflow; the headline quantitative claim is partly confounded by library retrieval, but the user-study and long-sequence results carry the paper.","tokens_in":18338,"tokens_out":867,"would_cite":true,"duration_ms":10339,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CustomDance claims that 3D dance generation is better done as a three-stage, human-in-the-loop authoring workflow—MLLM anchor planning, multimodal phrase retrieval, and diffusion-based gap filling and repair—than as a single end-to-end…","keywords":["3D dance generation","human-in-the-loop authoring","multimodal retrieval","diffusion inpainting","music-conditioned motion synthesis","multimodal large language model","choreography authoring","dance quality evaluation"],"falsifier":"Drop the roughly 3 hours of PromptHMR-reconstructed video and the StableMotion augmentations from the library, retrain the retriever and diffusion model, and re-run the Study 2 metrics on the same five 32-second clips; if FID_k and Div_k stay near 22.55 and 7.02 the reconstructed portion is not load-bearing, and if they jump toward MEGADance's 32.59 the headline quality numbers rest on data the paper never verifies against ground-truth motion. A second check: re-run the retriever on held-out phrases with the text modality masked off and compare Recall@10 to the reported 72.97, which isolates whether the claimed multimodal conditioning actually contributes.","tokens_in":17245,"feed_emoji":"🕺","tokens_out":11325,"duration_ms":90647,"temperature":0.7,"pith_summary":"CustomDance is an interactive system for authoring 3D dances, built around the claim that choreography is best treated as a human-in-the-loop process rather than a one-shot music-to-motion mapping. The paper proposes a three-stage workflow that mirrors how choreographers actually work: a multimodal language model reads the music and a global text prompt to propose where phrases should go and what should happen there; a retriever ranks short motion clips by local music, text, and body-part controls; and a diffusion model sews the chosen clips together and repairs flagged glitches. The system's evidence is that non-expert users can author a 32-second dance in about twelve minutes with no leftover kinematic artifacts, and that blinded raters prefer the results over two recent generative baselines in the large majority of comparisons. If the claim holds, the practical effect is that AI assistance for dance moves from generating something plausible to helping the user make what they actually want, with the human deciding structure and style at each step.","feed_headline":"Three-stage dance authoring tool outranks AI choreography baselines","feed_subtitle":"Planning, clip retrieval, and diffusion repair help users compose 3D dances faster, with fewer artifacts.","key_machinery":"The argument runs on four engineered objects working together. The anchor planner is an MLLM prompted to output JSON anchor–cue pairs that respect musical structure and a minimum temporal separation, converting abstract intent into concrete 4-second slots. The retriever is the search instrument: a contrastive trimodal model in which concatenated text and UI-control embeddings gate the music embedding through FiLM layers, trained with InfoNCE loss and modality dropout at 0.2 so missing inputs degrade gracefully, with library embeddings precomputed offline so each query costs about 0.7 seconds. The generative workhorse is a DDIM diffusion model trained on 8-second clips with an x0-prediction loss plus kinematic auxiliary losses, and its masked-denosing update—keeping known frames fixed while resynthesizing masked ones—implements both gap completion and joint-level repair without additional training. Finally, the diagnoser computes lightweight kinetic-energy proxies, absolute kinetic energy from linear joint velocities and relative kinetic energy from SO(3) angular velocities via the logarithm map, across six joint groups, so users can see exactly where root teleportation or limb twists occur and repair only those spans.","core_discovery":"The paper's central claim is that high-quality customized dance generation is an authoring problem, not just a synthesis problem, and that the right unit of design is a three-stage, coarse-to-fine pipeline. In the first stage, an MLLM (Gemini) analyzes the global music and a high-level text description and returns temporal anchors with imperative creative cues, which become 4-second phrase slots on a timeline. In the second, a contrastively trained trimodal retriever—jointly embedding local music (Librosa features), local text (CLIP), and six joint-group intensity/variety controls—ranks candidate phrases from a roughly 10.7-hour curated library (7.7 hours of FineDance motion capture, about 3 hours of PromptHMR-reconstructed video, augmented with StableMotion) and returns a top-10 list the user can preview and accept. In the third, a DDIM music-conditioned diffusion inpainter, built on a BiMamba–Transformer backbone, fills unassigned gaps by masked denoising and repairs user-selected intervals flagged by kinetic-energy visualizations across six joint groups, iterating until the user is satisfied. The paper reports that this workflow cuts average authoring time from 28.41 to 11.66 minutes and clip replacements from 25.16 to 5.32 versus a timeline-only editor, eliminates detector-flagged artifacts, and achieves FID 22.55, diversity 7.02/6.44, and BAS 0.233, beating Lodge, MEGADance, and the baseline editor on all objective metrics and earning 82–93% rank-1 preference rates in a blinded comparison.","pith_inferences":["My inference: the biggest measured win is search efficiency, not synthesis quality—authoring time dropped 59% when ranking replaced random browsing (28.28 to 11.66 minutes), so the retriever's ranking likely carries more of the user-facing value than the diffusion repair; a direct test would replace the learned retriever with music-similarity-only ranking and measure the drop.","The fixed 4-second slot follows eight-count phrasing at roughly 120 BPM, and nothing in the system adapts slot length to the music's tempo; a testable extension would make the slot duration tempo-aware and measure whether phrase appropriateness rises for slow or fast tracks.","Because final dances are mostly assembled from retrieved clips, the reported diversity partly measures the library's phrase coverage, which suggests the cheapest path to more diverse output is adding style-balanced motion capture rather than changing the generator.","The plan-retrieve-inpaint decomposition is not dance-specific; it could be tested as a general authoring paradigm for any timeline-based embodied content, such as gesture for virtual agents, martial-arts sequences, or character animation, where the user holds local creative intent."],"forward_implications":["Non-dancers can author a complete 32-second choreography in about 12 minutes with zero detector-flagged kinematic artifacts, down from 28+ minutes and more than five artifacts with a plain timeline editor.","For 90-second pieces, active authoring with CustomDance takes about 31 minutes, roughly half the time of professional scratch choreography, while expert raters judge the result close to the professionals on pose satisfaction and description fulfillment.","Blinded raters rank CustomDance dances first 82–93% of the time across music alignment, description fulfillment, and overall performance when compared against Lodge and MEGADance, so the authoring paradigm, not just the generator, is what users prefer.","The objective metrics (FID 22.55, diversity 7.02/6.44, BAS 0.233) suggest the composed dances are more realistic, more varied, and better beat-aligned than end-to-end baselines, which would make retrieval-plus-inpainting a competitive recipe for music-conditioned motion."],"supporting_citations":[{"why":"Supplies the 7.7-hour FineDance motion-capture corpus, the largest and most load-bearing portion of the dance library that both retrieval and diffusion training depend on.","marker":"[Li et al. 2023]"},{"why":"Reconstructs the roughly 3 hours of internet dance video into SMPL motion, the second library source whose alignment with FineDance the pipeline relies on.","marker":"[Wang et al. 2025]"},{"why":"Augments the combined corpus to produce the 4-second phrases and 8-second training clips used by both the retriever and the diffusion model.","marker":"[Mu et al. 2025]"},{"why":"Provides the editable music-to-dance diffusion formulation and the auxiliary kinematic losses that the Completer and Remaker backbone rest on.","marker":"[Tseng et al. 2023]"},{"why":"Supplies the BiMamba-Transformer hybrid architecture the generator follows and serves as a primary generative baseline the system must beat.","marker":"[Yang et al. 2025b]"},{"why":"Is the coarse-to-fine diffusion baseline for long dance generation and the source of the BAS metric and FSR evaluation used in the long-sequence study.","marker":"[Li et al. 2024a]"},{"why":"Is the MLLM that performs Stage 1 anchor-and-cue planning, turning music plus a global text prompt into temporal slots and creative cues.","marker":"[Team et al. 2023]"},{"why":"Maps local text prompts and default cues into the 512-dimensional vectors that condition the retriever's query.","marker":"[Radford et al. 2021]"},{"why":"Provides the audio features (MFCCs, CQT chroma, tempogram, onset/beat) that encode each local music segment for retrieval.","marker":"[McFee et al. 2015]"},{"why":"Modulates the music embedding with the concatenated text and UI-control vector inside the retriever's query encoder.","marker":"[Perez et al. 2018]"}],"fun_headline_variants":["CustomDance: three-stage AI choreography with user control","Interactive 3D dance generation via MLLM, retrieval, and diffusion","Coarse-to-fine dance authoring with multimodal control beats baselines","CustomDance: MLLM, retrieval, diffusion for user-refined 3D dance","Three-stage dance pipeline beats AI baselines in user study"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole system—the retrieved phrases and the diffusion model alike—rests on the curated 10.7-hour dance library being clean, consistently aligned, and representative of the styles users request; if noise from the reconstructed video, coordinate mismatches, or augmentation artifacts seep in, both the authoring experience and the objective metrics change regardless of how good the interaction design is.","fun_headline_variants_meta":{"raw":{"variants":["CustomDance: three-stage AI choreography with user control","Interactive 3D dance generation via MLLM, retrieval, and diffusion","Coarse-to-fine dance authoring with multimodal control beats baselines","CustomDance: MLLM, retrieval, diffusion for user-refined 3D dance","Three-stage dance pipeline beats AI baselines in user study"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000825,"raw_usage":{"total_tokens":3721,"prompt_tokens":1170,"completion_tokens":2551,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":786,"completion_tokens_details":{"reasoning_tokens":2454}},"tokens_in":786,"tokens_out":2551,"duration_ms":16910,"temperature":1.0,"reasoning_tokens":2454,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:51:39.714756+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Drop the roughly 3 hours of PromptHMR-reconstructed video and the StableMotion augmentations from the library, retrain the retriever and diffusion model, and re-run the Study 2 metrics on the same five 32-second clips; if FID_k and Div_k stay near 22.55 and 7.02 the reconstructed portion is not load-bearing, and if they jump toward MEGADance's 32.59 the headline quality numbers rest on data the paper never verifies against ground-truth motion. A second check: re-run the retriever on held-out phrases with the text modality masked off and compare Recall@10 to the reported 72.97, which isolates whether the claimed multimodal conditioning actually contributes.","supporting_citations":[],"review_version":1}