{"id":"f0d33cc1-09f6-41e4-bbd5-9512207952c9","arxiv_id":"2508.16911","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"MDD is the first dataset to pair text, music, and 3D duet dance motion, enabling two new text-conditioned duet generation tasks.","lead":"This paper introduces MDD, a new 10.3-hour motion capture dataset of professional duet dancing across 15 genres, synchronized with music and annotated with over 10,000 text descriptions. It defines two new tasks for generating duet dance from text and music, and evaluates baseline models on them.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unvalidated LLM refinement of text annotations risks breaking temporal/semantic alignment with motion and music, threatening the dataset's core claim of fine-grained text control.","rationale":"The reader's weakest assumption correctly identifies the LLM-based annotation refinement as the most fragile link in the central claim. I considered other potential concerns—the lack of public dataset access, the invisibility of baseline results in the provided text, and the subtle discrepancy between frame count and total duration—but none strike at the core argument as directly. Without verifiable alignment between the refined text and the motion/music, the dataset's 'fine-grained text control' is unsubstantiated; this is a necessary condition for both proposed tasks to be meaningful. The concrete test I propose (expert temporal segmentation of descriptions against motion) would directly measure whether the LLM introduces hallucinations or reorderings, and would settle the concern. I therefore agree with the reader's assessment and find no reason to change the CONDITIONAL verdict: the dataset's potential is clear, but validation of annotation fidelity is required before the central claim can be fully accepted.","tokens_in":8179,"tokens_out":3013,"duration_ms":36510,"concrete_test":"Stratify by genre and randomly select 100 annotated clips. For each clip, have two independent dance experts watch the motion with music and segment the LLM-processed description into temporally ordered moves, assigning each described move a time interval. Then compute (1) the percentage of described moves that are visually confirmed and correctly labeled, and (2) the percentage of clips where the order of described moves matches the temporal order in the motion. If move-label accuracy is below 90% or order errors occur in more than 5% of clips, the LLM refinement introduces alignment errors that would corrupt text-conditioned training. Compare these rates against the raw annotator descriptions to isolate the effect of LLM processing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—MDD is the first dataset integrating text, music, and duet dance with fine-grained control—depends critically on the text annotations being both semantically correct and temporally aligned with the paired motion and music. Section 3.2's 'Annotation Refining' step uses GPT-4o to rewrite raw annotator descriptions, as shown in Table 5. However, the paper reports no human validation of the refined descriptions: no inter-annotator agreement, no expert review, and no check that the LLM preserved the order and timing of described moves. The Table 5 examples show polished prose but do not demonstrate that the 'LLM-processed move names' correspond exactly to the motion, nor that temporal connectives like 'after which' match the actual sequence. If the LLM hallucinates a move, reorders events, or introduces details absent from the motion, then the proposed Text-to-Duet and Text-to-Dance Accompaniment tasks are conditioned on incorrect or misaligned descriptions. This directly undermines the dataset's value proposition and the validity of any baseline trained on it. The assumption is load-bearing: without text–motion alignment fidelity, the 'first text-and-music duet dataset' claim is hollow, regardless of the motion capture quality or scale.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MDD (Multimodal DuetDance), a dataset of 620 minutes (10.34 hours) of professional duet dance motion capture at 120 fps across 15 genres, synchronized with music and accompanied by more than 10K natural-language descriptions. The dataset supports two newly proposed tasks: Text-to-Duet, which generates both dancers' motion from music and a text prompt, and Text-to-Dance Accompaniment, which generates the follower's motion from the leader's motion, music, and text. The authors report statistics for genre coverage, dancer demographics, annotation counts, and BPM ranges, and they present baseline evaluations of an adapted interaction-generation method on the proposed tasks.","tokens_in":8448,"tokens_out":4843,"duration_ms":63089,"significance":"If the dataset and its annotations are reliable, MDD would be the first duet-dance dataset to combine motion, music, and fine-grained text, and it is substantially larger than existing duet-dance datasets (Duolando/DD100 at 1.95 hours and InterDance at 3.93 hours). The collection effort is a clear strength: 30 professional dancers, 15 genres across Ballroom, Latin, and Social styles, 120 fps OptiTrack capture, and explicit attention to marker-swear and post-processing. The two proposed tasks fill a genuine gap: existing duet-dance benchmarks lack text conditioning, and existing interactive-text benchmarks lack music and dance-specific vocabulary. The paper's central risk is annotation fidelity: the text annotations are the enabling modality, yet their semantic and temporal correctness is not validated. The baselines, while not the main contribution, give future work a starting point. The manuscript would be a community resource if the annotation-validation gap is closed and the dataset is released with clear documentation.","major_comments":[{"comment":"The central claim that MDD provides fine-grained text control depends on the correctness and temporal alignment of the GPT-4o-refined descriptions. The paper reports no human validation of the refined text: no inter-annotator agreement, no expert review, and no check that move names or temporal connectives ('after which', 'leading into') match the actual motion sequence. Table 5 shows polished output but does not demonstrate that the LLM did not hallucinate a move, reorder events, or add details absent from the motion. This is load-bearing because Text-to-Duet and Text-to-Dance Accompaniment condition on these descriptions. The manuscript must add a validation protocol, e.g., expert agreement rates, a hallucination audit, and a temporal-ordering check, or explicitly provide evidence that refinement preserves ground-truth alignment.","section":"§3.2 (Annotation Refining), Table 5"},{"comment":"Even before LLM refinement, no annotation-quality statistics are reported for the raw human descriptions. The paper states that annotators had diverse dance backgrounds but does not give the annotation instructions, number of annotators per clip, or any measure of agreement. Without inter-annotator agreement or a qualitative error analysis, the claim that the dataset contains >10K 'fine-grained' descriptions is an assertion rather than a demonstrated property. Please add annotation-protocol details and quantitative quality metrics for both raw and refined annotations.","section":"§3.2 (Data Collection) and Figure 10"},{"comment":"The annotation granularity is not specified: are text descriptions aligned to the entire clip, to fixed windows, or to time-stamped move segments? The examples in Figure 2 and Table 5 describe multi-step sequences with temporal ordering ('after which', 'leading into'), but no timestamps or segment boundaries are provided. For the proposed text-conditioned generation tasks, clip-level descriptions would give only coarse semantic control and would undercut the 'fine-grained' claim. The paper should state the temporal granularity of annotations and, if appropriate, provide segment-level annotations or timestamps.","section":"§3.1 and Figure 2"}],"minor_comments":[{"comment":"The column 'LLM-processed Move Name' is not defined. Clarify whether it is a free-form LLM output or a constrained label from a fixed vocabulary, and how it is extracted from the description.","section":"Table 5"},{"comment":"The claim that samples show 'high motion quality with rich annotations' is subjective. Please replace or supplement with quantitative indicators, e.g., marker-occlusion statistics, joint-angle smoothness, or a comparison of motion distributions across genres.","section":"Figure 2(a)"},{"comment":"Minor consistency check: 4.4M frames at 120 fps is approximately 10.19 hours, which is close to but not exactly '10.34 hours.' Either reconcile the numbers or clarify what the 620-minute figure includes (e.g., post-processed vs. raw capture time).","section":"Abstract and §3.1"},{"comment":"The paper should provide a data-availability statement with a direct download link, license, and usage terms. Also, the legal rationale for using copyrighted music excerpts under fair use is stated too briefly; given the dataset is to be distributed, provide more detail on the status of each audio track.","section":"References and §3.2.1"},{"comment":"There are several typos and spacing issues, e.g., 'with over10K' in the abstract and 'fro controlled release' in Table 5. A careful proofreading pass is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The dataset is a potentially valuable contribution, but the annotation-validation gap is the main barrier to acceptance. If the authors can supply even a modest human-evaluation study of the refined descriptions (e.g., expert agreement on move-name correctness and temporal order, plus a hallucination rate), the central claim would be substantially strengthened. The paper's 'first' claim is defensible only in the narrow duet-dance-with-text sense; the authors should make sure the comparison to TM2D and Inter-X is precise in the final version. I would support acceptance after these revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: MDD fills a real gap—no prior duet dance dataset has text, music, and motion in one package. At 10.34 hours across 15 genres with 30 trained dancers and over 10K descriptions, it's a substantial resource. The two proposed tasks (Text-to-Duet, Text-to-Dance Accompaniment) are natural and useful.\n\nThe data collection pipeline is described carefully: OptiTrack setup, post-processing, and the annotation tool. The genre and BPM distributions suggest good coverage. I believe the paper's core claim about being the first integrated text+music+duet dataset holds up against the cited literature.\n\nThe soft spot is annotation quality—specifically the GPT-4o refinement step in Section 3.2. The paper shows polished examples but gives no inter-annotator agreement, no checks that the LLM preserved the order and timing of moves, and no evidence that the refined descriptions match the motion. That's a load-bearing issue: if the text says 'after which' but the motion doesn't follow, the tasks are conditioned on noise. This is fixable—sample a subset and have experts verify—but the paper as written leaves it unaddressed. The reader's concern about missing baseline evaluations is also fair; the provided text doesn't show those experiments, though they may be in the accepted ICCV version.\n\nOne minor point: the music includes short excerpts of copyrighted songs under fair use. That could complicate future redistribution of the aligned audio. Worth clarifying.\n\nOverall, this is a solid dataset paper with a clear gap-filling contribution. It deserves serious peer review. If I were the reviewer, I'd ask for an annotation-quality analysis and dataset/code release before relying on it for fine-grained text control.","headline":"A genuinely new duet-dance dataset with text+music annotations, but the LLM-based annotation pipeline needs empirical validation before the fine-grained text control claim can be fully trusted.","tokens_in":8946,"tokens_out":2145,"would_cite":true,"duration_ms":25457,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces MDD, the first dataset to align duet dance motion, music, and fine-grained text, enabling text-controlled duet dance generation.","keywords":["duet dance generation","text-to-motion","music-conditioned generation","motion capture dataset","multimodal benchmark","lead-follower interaction","fine-grained text annotation"],"falsifier":"Take a random sample of MDD clips, have professional dancers identify every named move that occurs and its timing, then compare against the processed text descriptions; if a substantial share of descriptions contain moves not present in the clip, the claimed fine-grained text alignment is falsified.","tokens_in":8096,"feed_emoji":"💃","tokens_out":5486,"duration_ms":65085,"temperature":0.7,"pith_summary":"This paper introduces MDD, a large-scale duet dance dataset that joins three modalities that had not been paired for partner dancing: professional motion capture, synchronized music, and fine-grained natural-language descriptions. The authors aim to make duet dance generation controllable in the same way solo text-to-motion is controllable, by giving models text that specifies moves, spatial relationships, and rhythm while music supplies timing. They claim MDD is the first such dataset, with 10.34 hours of motion capture across 15 genres and more than 10,000 text descriptions, and they define two new tasks around it. A reader should care because the dataset turns an under-explored problem—two dancers coordinating as leader and follower—into a benchmarkable generation task.","feed_headline":"First dataset links text, music, and duet dance motion","feed_subtitle":"10,000+ descriptions across 15 genres enable two new text-and-music dance tasks.","key_machinery":"The carrying mechanism is the aligned triple of motion capture clips, music, and fine-grained text annotations. Two design choices do the work: genre-balanced sampling across 15 Latin, Ballroom, and Social dance genres, and a two-stage annotation pipeline in which human annotators produce raw descriptions and a language model refines them into polished move-name-plus-description text. The annotation vocabulary—spatial relationships, body movement, and rhythm—is what lets text control both leader and follower roles.","core_discovery":"The central claim is that duet dance can be captured as a three-way aligned multimodal benchmark: over 4.4 million frames of 120 fps motion capture from 30 dancers, music synchronized to the recordings, and text annotations that name specific moves and describe spatial relationships, body movements, and rhythm. The paper argues this is the first dataset to integrate human motion, music, and text for duet dance, going beyond prior duet dance datasets that pair motion with music but lack text, and beyond two-person interaction datasets that lack dance-specific movement vocabulary and music. On this foundation the paper defines two tasks: Text-to-Duet, where both dancers' motions are generated","pith_inferences":["Not claimed in the paper: the same annotation taxonomy could transfer to group choreography or contact sports, where spatial relationships and rhythmic cues also matter.","Not claimed in the paper: a testable extension would be to have expert dancers verify whether every named move in the processed text actually occurs in the corresponding motion clip, quantifying annotation precision beyond the paper's illustrative examples.","Not claimed in the paper: the two tasks could feed interactive applications where a user supplies a text prompt and a virtual partner dances in real time; the paper itself stops at static generation benchmarks.","Not claimed in the paper: whether text descriptions actually improve over music-only conditioning for duet dance is not settled by the baseline experiments; a controlled ablation with and without text would isolate that contribution."],"forward_implications":["Text-to-Duet becomes a well-defined benchmark task with shared data, allowing models to be compared on leader-follower coordination rather than on bespoke settings.","The 15-genre, wide-BPM coverage makes it possible to test whether a single text-conditioned model transfers across dance styles.","Text-to-Dance Accompaniment gives a concrete setting for reactive motion synthesis: the follower's motion can be evaluated relative to the leader, the music, and the text, not just for visual plausibility.","Because the annotations name specific moves, downstream work can use them as pseudo-labels for move-level retrieval and choreography parsing."],"supporting_citations":[{"why":"Prior duet dance dataset pairing motion with music but lacking text annotations; defines the gap MDD fills.","marker":"[31]"},{"why":"Other prior duet dance dataset with 3.93 hours across 15 genres; provides the size and coverage comparison.","marker":"[17]"},{"why":"Supplies the two-person interaction generation framework and the baseline model adapted for the Text-to-Duet task.","marker":"[20]"},{"why":"Text-annotated two-person interaction dataset that motivates fine-grained text conditioning for interactive motion.","marker":"[38]"},{"why":"Standard music-dance benchmark defining the solo dance generation task against which MDD's duet contribution is positioned.","marker":"[14]"},{"why":"Largest solo dance motion capture dataset, used as a reference for high-quality data and genre coverage.","marker":"[15]"},{"why":"Language model used in the annotation refining step to produce the processed text descriptions.","marker":"[11]"},{"why":"Spherical interpolation method used to fill missing marker data during motion post-processing.","marker":"[32]"}],"fun_headline_variants":["New dataset pairs 10K text notes with duet dance motion","Duet dance dataset syncs motion, music, and 10K text prompts","First text-and-music benchmark for generating duet dance","620 minutes of duet motion plus text for dance AI","Two tasks debut: generate duet dance from text and music"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The central assumption is that the automated cleanup of the written descriptions keeps them accurately tied to what the two dancers actually do in the motion capture; if that tie is broken, the text control the dataset promises is not really there.","fun_headline_variants_meta":{"raw":{"variants":["New dataset pairs 10K text notes with duet dance motion","Duet dance dataset syncs motion, music, and 10K text prompts","First text-and-music benchmark for generating duet dance","620 minutes of duet motion plus text for dance AI","Two tasks debut: generate duet dance from text and music"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00011,"raw_usage":{"total_tokens":861,"prompt_tokens":688,"completion_tokens":173,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":432,"completion_tokens_details":{"reasoning_tokens":84}},"tokens_in":432,"tokens_out":173,"duration_ms":2696,"temperature":1.0,"reasoning_tokens":84,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:07:16.925060+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of MDD clips, have professional dancers identify every named move that occurs and its timing, then compare against the processed text descriptions; if a substantial share of descriptions contain moves not present in the clip, the claimed fine-grained text alignment is falsified.","supporting_citations":[{"cited_title":"Intergen: Diffusion-based multi-human motion genera- tion under complex interactions","cited_arxiv_id":null,"evidence_quote":"Supplies the two-person interaction generation framework and the baseline model adapted for the Text-to-Duet task."},{"cited_title":"Inter-x: Towards versatile human- human interaction analysis","cited_arxiv_id":null,"evidence_quote":"Text-annotated two-person interaction dataset that motivates fine-grained text conditioning for interactive motion."},{"cited_title":"Ai choreographer: Music conditioned 3d dance generation with aist++","cited_arxiv_id":null,"evidence_quote":"Standard music-dance benchmark defining the solo dance generation task against which MDD's duet contribution is positioned."},{"cited_title":"Finedance: A fine-grained choreography dataset for 3d full body dance generation","cited_arxiv_id":null,"evidence_quote":"Largest solo dance motion capture dataset, used as a reference for high-quality data and genre coverage."},{"cited_title":"Gpt-4o: The cutting-edge advancement in multimodal llm","cited_arxiv_id":null,"evidence_quote":"Language model used in the annotation refining step to produce the processed text descriptions."},{"cited_title":"Smith and J","cited_arxiv_id":null,"evidence_quote":"Spherical interpolation method used to fill missing marker data during motion post-processing."}],"review_version":1}