{"id":"056ba4d1-3421-4582-9023-7da8a498abb1","arxiv_id":"2608.07993","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"MRBench is a multi-source, balanced, multi-granular human motion-text retrieval benchmark, and the proposed granularity-aware adapters improve mixed-granularity retrieval without degrading standard-caption retrieval.","lead":"This paper introduces MRBench, a motion-text retrieval benchmark with 3,390 motions from mocap, video, and generative sources, paired with 10,170 captions at three granularities. It also proposes a granularity-aware retrieval model that improves fine-grained retrieval while preserving standard-caption performance.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Benchmark validity hinges on unquantified LLM/VLM curation; without thresholds, pass rates, or inter-annotator agreement, misaligned or hallucinated captions would propagate into all retrieval results and the model comparison.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing gap: the benchmark's reliability depends on an LLM/VLM pipeline whose acceptance thresholds, pass rates, and human-validation statistics are unreported. I agree with that assessment, and I would not move the verdict: the paper is plausibly correct and the issues are addressable, but the current evidence does not fully establish the central reliability claim. I also considered two other candidate concerns. First, the duplicate Ours/MoPatch row in Table 1's Standard block is a real data-integrity flag but is cosmetic relative to the curation issue; it can be fixed by clarifying that the frozen anchor is identical for standard queries. Second, training on MRBench-Train and evaluating on MRBench is a same-distribution split rather than true cross-dataset transfer, but the paper labels it as a general-purpose testbed and does not overclaim this particular comparison. Neither is as load-bearing as the unquantified curation pipeline. The proposed concrete test—human double-annotation on a stratified sample plus a Gemini-versus-human verdict comparison—would settle whether the captions are actually motion-verified and would also reveal whether fine-grained rewriting introduces hallucinated details. Until such evidence is provided, the conditional verdict remains the right call.","tokens_in":14591,"tokens_out":4243,"duration_ms":44263,"concrete_test":"Release a stratified random sample of 300 MRBench motions spanning all four sources and all 118 categories in proportion, and have at least two independent human annotators judge, for each of the three captions, whether every described action is observable in the rendered motion. Report per-granularity acceptance rates and Cohen's kappa. Additionally, on a pre-rewrite sample, compare Gemini SSAE verdicts with human verdicts on the same motion-text pairs. If fine-grained acceptance falls below 95% or kappa below 0.7, the 'motion-verified, semantically unambiguous' claim is not established and the reported evaluation numbers require re-examination.","verdict_should_be":"UNCHANGED","load_bearing_attack":"MRBench's central claim—that it contains 'motion-verified, well-aligned' captions at three granularities—rests entirely on the multi-stage pipeline described in the 'MRBench for Motion-Text Retrieval' section: GPT-5.5 scoring and filtering, Qwen taxonomy assignment, Gemini-3-pro-preview SSAE verification on 5 FPS frames, and Gemini multi-granular rewriting, followed by an unquantified manual validation step. The paper reports no acceptance thresholds for the GPT-5.5 scores, no pass rates or rationale-quality checks for the Gemini SSAE step, no per-granularity verification of the rewritten concise and fine-grained captions, and no inter-annotator agreement or error analysis for the manual review. This gap is load-bearing because every headline result—the cross-dataset generalization gap, the granularity sensitivity, and the proposed model's fine-grained gains—is computed against these captions. If Gemini's binary 'is the primary motion visible?' check misses fast or subtle motions at 5 FPS, or if the rewriting step hallucinates body-part or temporal details, then the fine-grained queries are not testing semantic granularity but rather LLM stylistic bias. The same Gemini model also both verifies and rewrites, so any systematic alignment bias is shared between the verification and the generated captions. Without quantification, the benchmark's reliability cannot be independently checked, which undermines the foundational assumption of the entire evaluation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MRBench, a motion-text retrieval benchmark containing 3,390 motion sequences from MoCap, in-the-wild video, synthetic video, and generative-model sources, with 118 fine-grained categories and three caption granularities (concise, standard, fine-grained), totaling 10,170 captions. The benchmark is built through a multi-stage pipeline: GPT-5.5-based caption filtering, Qwen-based taxonomy assignment, balanced sampling, Gemini-based semantic alignment verification, Gemini-based multi-granular rewriting, and manual validation. The paper also proposes a granularity-aware retrieval model that freezes a standard-caption-aligned dual encoder and trains lightweight motion extractors and text adapters on LLM-rewritten concise and fine-grained captions, with score fusion at inference. The experiments evaluate several baselines on MRBench under single-granularity and mixed-granularity protocols, reporting cross-dataset generalization gaps and granularity sensitivity, and the proposed model is shown to improve fine-grained retrieval while preserving standard-caption performance.","tokens_in":15059,"tokens_out":11057,"duration_ms":114769,"significance":"If the benchmark's curation claims are valid, MRBench addresses a real gap in motion-text retrieval evaluation, which currently relies on homogeneous, imbalanced datasets with repetitive captions. The statistical analysis of HumanML3D and KIT-ML in Fig. 2 is useful and gives concrete evidence of the ambiguity problems in existing benchmarks. The multi-source composition and the three-granularity annotation scheme are sensible design choices, and the proposed model, which anchors on a frozen standard-aligned encoder, is simple and well motivated. The cross-dataset generalization results are falsifiable and potentially valuable for the community. However, the benchmark's foundational reliability, the proposed model's gains, and the claimed score comparability in mixed-granularity retrieval cannot currently be assessed because the curation pipeline lacks quantified acceptance criteria, the evaluation may be confounded with LLM caption style, and no data or code release is stated.","major_comments":[{"comment":"The central claim that MRBench contains 'motion-verified, well-aligned' captions is not independently checkable because no quantitative acceptance criteria are reported for the curation stages. 'Text-Guided Candidate Filtering' lists GPT-5.5 scores (visual noise, object noise, motion specificity, low ambiguity) but gives neither thresholds nor score distributions; 'Semantic Alignment and Disambiguation' reports no SSAE pass rate, no alignment acceptance threshold, and no rationale-quality check for the Gemini-3-pro-preview verdicts; and the manual review in 'Multi-granular Text Expansion' provides no inter-annotator agreement, correction statistics, or per-granularity verification counts. Because Tables 1-3 evaluate retrieval on these captions, the validity of every headline result is currently unverified.","section":"Benchmark Construction"},{"comment":"The fine-grained evaluation is exposed to a style-circularity risk. The MRBench fine-grained captions are produced by Gemini-3-pro-preview from standard captions plus alignment rationale, while the proposed model's concise and fine-grained branches are trained with pseudo-labels generated by Qwen2.5-7B-Instruct from the same standard captions. Reported fine-grained gains may reflect adaptation to LLM paraphrase statistics rather than genuine motion understanding. This is a testable concern: report results on held-out human-written fine-grained queries or on LLM-generated captions with style perturbed, and report how model performance changes with caption style statistics. Note also that Gemini is used for both SSAE verification and rewriting, so any systematic bias in its motion-visibility judgments is shared between the verification and the generated captions.","section":"Granularity-Aware Retrieval Model (Training and Inference)"},{"comment":"The statement that Eq. (7) 'makes fused scores across different textual granularities directly comparable' is stronger than the formula supports. Dividing by (1+alpha_g) applies a fixed per-granularity scale but does not correct for different location or shape of the global similarity distributions across concise, standard, and fine-grained queries. In the Mixed3 protocol this may leave per-granularity offsets in the ranking. Please report score distributions by granularity on the validation split, add an explicit per-granularity calibration if needed, and demonstrate that Mixed3 rankings are robust to monotone per-granularity transformations.","section":"Granularity-Aware Inference, Eq. (7)"},{"comment":"There is a contradiction in the training-data description. The 'Benchmark Statistics' section says that MRBench-Train 'follows the same annotation protocol' as MRBench, which includes multi-granular rewriting, but the method section says that concise and fine-grained variants are expanded by Qwen2.5-7B-Instruct because the training set consists of standard-caption pairs only. Please clarify whether MRBench-Train contains multi-granular captions or only standard captions, and specify exactly what 'same annotation protocol' means.","section":"Benchmark Statistics and Granularity-Aware Retrieval Model"},{"comment":"The paper provides no data or code availability statement, release URL, or licensing information for MRBench. A benchmark paper's central artifact must be publicly available for the retrieval scores in Tables 1-3 to be reproducible and for the claimed community testbed role to be realized; please add a clear availability statement.","section":"Benchmark Statistics"}],"minor_comments":[{"comment":"In Table 1, where all models are trained on H3D, the proposed model is absent from the H3D Standard block; since the model is reported to preserve standard-caption performance, please include this entry or state explicitly why it is omitted.","section":"Table 1"},{"comment":"In Table 2, the MotionMillion block does not include SGAR although SGAR is listed as a baseline; state whether the results are unavailable or were omitted.","section":"Table 2"},{"comment":"The label 'Motion Diverity' in Figure 3 should be corrected to 'Motion Diversity'.","section":"Figure 3"},{"comment":"In Eq. (2), specify the softmax axis and state the dimensions of Wk and Wv so that the attention pooling over the motion token sequence is unambiguous.","section":"Eq. (2)"},{"comment":"Figure 2(c) references 'mirrored pairs' without a definition; please explain how mirrored pairs are identified and merged for each dataset, since the merging rule affects the reported ambiguity percentages.","section":"Figure 2(c)"}],"recommendation":"major_revision","confidential_remarks":"The benchmark idea is potentially valuable and the empirical claims are worth pursuing, but the current manuscript leaves the curation pipeline unquantified and the proposed model's fine-grained gains potentially confounded with LLM style. These are fixable with additional transparency: report thresholds, pass rates, inter-annotator agreement, score distributions, and held-out human-written queries, and release the data and code. I would not reject on the style-circularity concern alone because it can be tested empirically. The main risk is that the benchmark will be published as a black-box artifact without the missing statistics."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"MRBench is a serious attempt at fixing real problems in HumanML3D and KIT-ML. The analysis of shared captions and retrieval ambiguity is solid, and the multi-source, balanced design is a clear step forward. The three-granularity annotation scheme is a genuine novelty, and the granularity-aware adapter method is reasonable engineering that does what it claims: preserving standard-caption performance while improving fine-grained retrieval.\n\nThe soft spots are about evidence, not ideas. The curation pipeline is plausible but opaque: no thresholds for the GPT-5.5 filtering, no pass rates or rationale checks for the Gemini SSAE verification, no inter-annotator agreement for the manual validation. The stress-test note is right that the benchmark's core claim—'motion-verified, well-aligned' captions—rests on these numbers. At 5 FPS, a binary VLM check can miss fast or subtle motions; and since Gemini both verifies and rewrites, systematic alignment bias would propagate. The absence of release is the biggest issue: a benchmark is only useful if others can use it. I would not desk-reject, but I would require data/code release and curation statistics in revision.\n\nOne correction to the reader's take: Table 1's Ours row on Standard is identical to MoPatch by design—the frozen anchor scores standard queries. On Concise the two rows are not identical (e.g., MedR 79 vs 82). So that is not a copy/paste error, though the paper should say explicitly that the standard row is the anchor's output.\n\nAlso, the circularity concern is real but milder than it looks: benchmark captions are Gemini-generated, training pseudo-labels are Qwen-generated, so the model is not directly fitting the eval caption style. Still, both are LLM styles, and the benchmark's discriminative value depends on how well the rewriting avoids hallucinated body-part details.\n\nOverall, a useful paper for the motion-language community. The benchmark should be reviewed seriously, but only accepted after the pipeline is quantified and the data released.","headline":"A genuinely useful benchmark proposal for motion-text retrieval, but the lack of curation transparency and missing data release keep it from being trustworthy yet.","tokens_in":15457,"tokens_out":2373,"would_cite":false,"duration_ms":23512,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces MRBench, a heterogeneous, multi-granularity benchmark for human motion-text retrieval, and a granularity-aware model that improves mixed-granularity retrieval without sacrificing standard-caption performance.","keywords":["human motion-text retrieval","benchmark","multi-granular captions","cross-modal alignment","granularity-aware retrieval","motion-language datasets","contrastive learning","dataset curation"],"falsifier":"Take a random sample of MRBench motions and have independent human annotators, who have not seen the construction pipeline, judge whether each concise, standard, and fine-grained caption matches the motion; if a substantial fraction of captions are judged mismatched or ambiguous, the claim that MRBench provides reliable, motion-verified, unambiguous captions is false. Alternatively, recompute the category distribution after merging semantically identical labels, and if it becomes as concentrated as HumanML3D's distribution, the balance claim fails.","tokens_in":1989,"feed_emoji":"🕺","tokens_out":3242,"duration_ms":106304,"temperature":0.7,"pith_summary":"The paper argues that existing motion-text retrieval benchmarks, such as HumanML3D and KIT-ML, largely measure how well models match repetitive indoor motion-capture captions rather than true cross-modal understanding. It reports that in those datasets many test queries share identical text with other gallery motions, so even correct retrievals are often counted as wrong. To address this, the authors construct MRBench, a benchmark of 3,390 motions drawn from motion capture, in-the-wild video, synthetic video, and generative models, covering 118 categories, with each motion labeled by concise, standard, and fine-grained captions for a total of 10,170 captions. They show that current retrieval models drop sharply on MRBench and are sensitive to caption granularity. They also propose a lightweight granularity-aware model that freezes a standard-caption-aligned backbone and adds granularity-specific extractors and adapters, which improves fine-grained and mixed-granularity retrieval without lowering standard-caption performance.","feed_headline":"New motion-text benchmark exposes cross-domain retrieval gaps","feed_subtitle":"3,390 motions from four sources, captioned at three detail levels, reveal current models overfit repetitive indoor text.","key_machinery":"The load-bearing object is the benchmark itself, MRBench, containing 3,390 motions standardized to a unified skeleton at 20 FPS with 10,170 captions. It is built by a four-stage pipeline: text-guided filtering using rule-based checks and large-language-model scoring; taxonomy-guided balanced sampling into 7 coarse, 33 middle-level, and 118 fine-grained categories with embedding-based deduplication; semantic alignment verification using a vision-language model that checks whether the primary motion described in a caption is visible in sampled frames; and multi-granular rewriting, where the same vision-language model generates concise and fine-grained variants that are then manually validated. The proposed model uses a frozen dual encoder trained on standard captions as an alignment anchor, then trains lightweight attention-pooling motion extractors and text-projection adapters on LLM-rewritten concise and fine-grained captions as pseudo-supervision. At inference, the known query granularity selects a branch, and for non-standard granularities the global and adapted cosine similarities are fused with a granularity-specific weight so that scores remain comparable across all description levels.","core_discovery":"The central claim is that a useful motion-text retrieval benchmark must vary both the motion distribution and the textual granularity to expose genuine alignment ability. The paper claims MRBench is the first such benchmark, combining heterogeneous motion sources with balanced category coverage and unique, motion-verified, multi-granular captions. On this benchmark, the authors observe a substantial cross-dataset generalization gap: methods trained on HumanML3D perform far worse on MRBench, and all evaluated baselines are sensitive to whether the query is concise, standard, or fine-grained. They further claim that their granularity-aware model, which keeps a frozen standard-caption-aligned dual encoder and learns granularity-specific attention-based motion extractors and text adapters with calibrated score fusion, improves fine-grained and mixed-granularity retrieval while preserving standard-caption retrieval exactly.","pith_inferences":["A natural next step is to publish inter-annotator agreement statistics from the manual validation pass; until then, the paper's claim that all captions are motion-verified cannot be independently checked.","The benchmark's usefulness depends on caption uniqueness, and the paper reports that after merging mirrored pairs, 19.6% and 5.9% of existing benchmark queries remain ambiguous while MRBench's rate is roughly zero; auditing that comparison directly would test the benchmark's discriminative advantage.","The granularity-aware split-branch recipe could transfer to image-text or video-text retrieval, where caption specificity also changes retrieval difficulty.","Because the concise and fine-grained pseudo-labels come from a large language model, part of the model's gain may reflect matching LLM rewriting style rather than deeper motion understanding; rewriting the same motions with a different model and measuring retention would separate those effects."],"forward_implications":["On MRBench, models trained on HumanML3D lose a large share of their recall, so in-domain scores on existing benchmarks overstate robustness to heterogeneous motion and text.","Retrieval quality depends strongly on query granularity: concise and fine-grained captions are harder than standard captions for all tested baselines.","Mixed-granularity retrieval requires score calibration across branches; removing that calibration degrades motion-to-text ranking substantially.","Training on MRBench-Train transfers better to MRBench than training on KIT-ML, HumanML3D, or MotionMillion, indicating that motion diversity and annotation discriminability matter more than raw data scale.","The granularity-aware model improves fine-grained and mixed-granularity retrieval without changing standard-caption results, demonstrating that robustness to non-standard queries can be added without sacrificing the original alignment."],"supporting_citations":[{"why":"Supplies the entire candidate motion-text pool and the taxonomy context that MRBench is built from.","marker":"ViMoGen-228K (Lin et al. 2025)"},{"why":"Provides the hierarchical taxonomy inspiration, the Structured Semantic Alignment Evaluation procedure, and synthesized motions used to supplement rare categories.","marker":"HY-Motion (Wen et al. 2025)"},{"why":"It is the dominant existing benchmark whose homogeneity and text repetition the paper measures and uses as a comparison test set.","marker":"HumanML3D (Guo et al. 2022)"},{"why":"It is the other standard benchmark used for the same motivation and for cross-dataset comparison.","marker":"KIT-ML (Plappert, Mandery, and Asfour 2016)"},{"why":"It is the frozen standard-caption-aligned anchor model that the proposed granularity-aware method builds on.","marker":"MoPatch (Yu, Tanaka, and Fujiwara 2024)"},{"why":"It is a structural-augmentation retrieval baseline whose scores are compared on MRBench.","marker":"SGAR (Zhang et al. 2025a)"},{"why":"It is the contrastive text-to-motion retrieval baseline evaluated in the main tables.","marker":"TMR (Petrovich, Black, and Varol 2023)"},{"why":"It performs taxonomy classification for remaining samples and rewrites concise and fine-grained captions used as pseudo-supervision.","marker":"Qwen2.5-7B-Instruct (Yang et al. 2025)"},{"why":"It verifies semantic alignment on sampled motion frames and generates the multi-granular caption rewrites.","marker":"Gemini-3-pro-preview (Google DeepMind 2025)"},{"why":"It scores caption discriminability and compactness during candidate filtering and generates motion scripts for rare categories.","marker":"GPT-5.5 (OpenAI 2026)"}],"fun_headline_variants":["MRBench: 3,390 motions, 3 granularities, big retrieval gaps","Benchmark reveals models fail on varied motion text queries","New motion-text benchmark stresses granularity and domain gaps","MRBench: heterogeneous motion data exposes retrieval weaknesses","Granularity-aware model beats baselines on new MRBench test"],"cache_read_input_tokens":17536,"weakest_assumption_plain":"The benchmark's reliability rests on the assumption that the automated filtering and verification pipeline, plus a small unquantified manual review, correctly judged every caption to match its motion and to be unambiguous.","fun_headline_variants_meta":{"raw":{"variants":["MRBench: 3,390 motions, 3 granularities, big retrieval gaps","Benchmark reveals models fail on varied motion text queries","New motion-text benchmark stresses granularity and domain gaps","MRBench: heterogeneous motion data exposes retrieval weaknesses","Granularity-aware model beats baselines on new MRBench test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000557,"raw_usage":{"total_tokens":2673,"prompt_tokens":993,"completion_tokens":1680,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":1595}},"tokens_in":609,"tokens_out":1680,"duration_ms":11106,"temperature":1.0,"reasoning_tokens":1595,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:33:49.687243+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of MRBench motions and have independent human annotators, who have not seen the construction pipeline, judge whether each concise, standard, and fine-grained caption matches the motion; if a substantial fraction of captions are judged mismatched or ambiguous, the claim that MRBench provides reliable, motion-verified, unambiguous captions is false. Alternatively, recompute the category distribution after merging semantically identical labels, and if it becomes as concentrated as HumanML3D's distribution, the balance claim fails.","supporting_citations":[],"review_version":1}