{"id":"3672a739-9f5b-4e38-a619-b26d87c54001","arxiv_id":"2605.29488","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"AnyMo is a masked-modeling framework for any-modality human motion generation trained on the new OmniHuMo dataset of 5,000+ hours of multimodal motion sequences.","lead":"The paper introduces OmniHuMo, a dataset of over 5,000 hours of motion sequences with aligned text, speech, music and trajectory annotations, plus AnyMo, a masked-modeling transformer that generates motion from arbitrary combinations of those modalities. A smart generalist might read it to see how large aligned datasets and unified architectures could expand controllable motion synthesis for animation and robotics.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"OmniHuMo alignment precision and cross-modal generalization without fine-tuning unverified","rationale":"Reader's weakest_assumption exactly matches the load-bearing precondition for the scaling claim. Full-text details (if present) would need to supply the missing alignment validation or ablations to overturn it; absent that evidence the UNVERDICTED stance is appropriate.","tokens_in":1630,"tokens_out":305,"duration_ms":15415,"concrete_test":"From the released OmniHuMo subset, randomly sample 500 sequences and compute modality-pair alignment scores (CLIP text-motion cosine, audio-motion beat alignment, trajectory-motion endpoint error); if mean alignment <0.75 or variance >0.2 across pairs, retrain AnyMo on the filtered subset and measure FID / user preference drop on held-out multi-modal prompts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that OmniHuMo's 5,000+ hours / 3.2M sequences supply precisely aligned multimodal annotations (text, speech, music, trajectory) sufficient for one masked-modeling transformer to generalize across arbitrary combinations without modality-specific fine-tuning. The abstract asserts this scale and precision but supplies no description of the alignment procedure, no quantitative alignment metrics, and no ablation showing performance parity (or improvement) when modalities are dropped or combined. If alignment noise or coverage gaps exist, the Residual FSQ tokenizer + masked transformer cannot reliably learn the claimed cross-modal interactions.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces OmniHuMo, a dataset of over 5,000 hours of motion across 3.2 million sequences with precisely aligned multimodal annotations (text, speech, music, trajectory), and AnyMo, a unified framework that combines a Residual FSQ motion tokenizer with a masked modeling transformer to support high-fidelity conditional motion synthesis under arbitrary modality combinations while providing flexible spatial and stylistic control.","tokens_in":1724,"tokens_out":335,"duration_ms":16489,"significance":"If the dataset alignment precision and cross-modal generalization claims hold, the work would be significant for addressing data scarcity in multimodal motion generation and for demonstrating scalable masked modeling across modality combinations, with potential impact on flexible control in computer vision and robotics applications.","major_comments":[{"comment":"Abstract: The central claim that OmniHuMo supplies 'precisely aligned multimodal annotations' at sufficient scale and quality for a single transformer to generalize across arbitrary modality combinations without modality-specific fine-tuning is load-bearing, yet the manuscript supplies no description of the alignment procedure, no quantitative alignment metrics, and no ablation results showing performance when individual modalities are dropped or combined.","section":"Abstract"},{"comment":"Abstract: The assertion of 'high-fidelity synthesis' and 'flexible control over both spatial and stylistic attributes' under arbitrary modality combinations lacks supporting cross-modal ablation tables or comparisons to modality-specific baselines, which are required to substantiate that the Residual FSQ tokenizer plus masked transformer reliably learns the claimed interactions.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback highlighting the need to better substantiate the alignment claims and cross-modal performance assertions. We will revise the manuscript to include the requested details and experiments.","responses":[{"response":"We agree that the alignment procedure and supporting metrics are essential to substantiate the central claims. While the manuscript describes the overall data curation process in Section 4, it does not provide a dedicated account of the alignment methodology or quantitative synchronization metrics. We will add a new subsection detailing the alignment procedure (including tools and protocols used for text-speech-music-trajectory synchronization) along with quantitative metrics such as temporal offset statistics and inter-annotator agreement scores. We will also include ablations that systematically drop or combine modalities to demonstrate generalization without modality-specific fine-tuning.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The central claim that OmniHuMo supplies 'precisely aligned multimodal annotations' at sufficient scale and quality for a single transformer to generalize across arbitrary modality combinations without modality-specific fine-tuning is load-bearing, yet the manuscript supplies no description of the alignment procedure, no quantitative alignment metrics, and no ablation results showing performance when individual modalities are dropped or combined."},{"response":"We acknowledge that the current experiments section would benefit from explicit cross-modal ablation tables and direct comparisons against modality-specific baselines to more rigorously support the high-fidelity and flexible control claims. We will add these tables, including quantitative results for all modality combinations versus single-modality baselines, along with qualitative examples demonstrating spatial and stylistic control. This will be incorporated into the revised experiments and supplementary material.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The assertion of 'high-fidelity synthesis' and 'flexible control over both spatial and stylistic attributes' under arbitrary modality combinations lacks supporting cross-modal ablation tables or comparisons to modality-specific baselines, which are required to substantiate that the Residual FSQ tokenizer plus masked transformer reliably learns the claimed interactions."}],"tokens_in":1268,"tokens_out":427,"duration_ms":14674,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's real addition is OmniHuMo, a 5,000-hour dataset with 3.2 million sequences carrying aligned text, speech, music, and trajectory labels, plus AnyMo, which combines a residual FSQ motion tokenizer with a masked transformer to accept arbitrary modality subsets in one model.\n\nThat scale and the unified setup are the parts worth noticing. Prior motion generators usually lock to one or two control signals and need separate heads or fine-tuning; trying to train one transformer that handles any mix without retraining is a reasonable next step and matches the direction in other multimodal generative work.\n\nThe soft spot is exactly the one the stress-test flags. The abstract and the central claim both depend on the dataset supplying clean, precise cross-modal alignment at that volume, yet the description of how alignment was done, what the quantitative checks were, and whether performance holds when modalities are dropped or swapped is thin. Without those ablations or alignment metrics, it is hard to know whether the reported high-fidelity results actually come from learned cross-modal interactions or from the model mostly using the strongest single signal. The tokenizer and masking choices are standard enough that they do not carry the load by themselves.\n\nThis is for groups already working on conditional human motion for animation or robotics who need larger training sets and are willing to test a unified architecture. Readers who care about scaling laws in motion models will want the dataset if it is released with the alignment code.\n\nSend it to peer review. The data contribution is large enough to justify referee time even if the generalization experiments need tightening.","headline":"AnyMo ships a genuinely large new motion dataset and a single masked transformer for variable modalities, but the alignment quality and no-fine-tune generalization rest on details that need checking.","tokens_in":2191,"tokens_out":400,"would_cite":false,"duration_ms":18256,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"AnyMo enables high-fidelity human motion generation from arbitrary combinations of text, speech, music, and trajectory controls using a single masked modeling transformer.","keywords":["human motion generation","multimodal conditioning","masked modeling","motion synthesis","transformer","conditional generation"],"falsifier":"A controlled test in which the model receives deliberately misaligned modality labels for a held-out combination and produces lower fidelity or less controllable outputs than single-modality baselines.","tokens_in":2545,"feed_emoji":"🚶","tokens_out":529,"duration_ms":19347,"temperature":0.7,"pith_summary":"The paper introduces a dataset of over 5,000 hours of motion sequences paired with aligned text, speech, music, and trajectory labels. It then builds a single model that tokenizes motion and applies masked modeling to accept any subset of those labels as conditioning. This setup removes the need for separate architectures per modality combination. If the approach holds, motion synthesis systems can accept mixed control signals without retraining or redesign for each new input type.","feed_headline":"One transformer generates motion from any mix of text, speech and paths","feed_subtitle":"Trained on 5,000 hours of aligned multimodal sequences, the model handles flexible spatial and stylistic controls without fixed modality set","key_machinery":"The masked modeling transformer that reconstructs motion tokens while conditioned on arbitrary modality inputs.","core_discovery":"AnyMo pairs a Residual FSQ-based motion tokenizer with a scalable masked modeling transformer trained on the OmniHuMo dataset of over 5,000 hours of multimodal motion sequences, enabling high-quality synthesis controlled by any subset of the available modalities.","pith_inferences":["The same masked modeling pattern could be tested on other sequential outputs such as full-body video or audio waveforms.","Further scaling of the dataset size would be expected to improve fidelity and control precision in the same way observed for language models."],"forward_implications":["A single model can accept mixed spatial controls such as trajectories together with stylistic controls such as text or music.","No separate fine-tuning or architecture changes are required when the set of available control signals changes.","Cross-modal interactions are learned directly through masked reconstruction on large aligned sequences."],"fun_headline_variants":["AnyMo transformer generates motion from any mix of text speech and paths","Masked modeling transformer enables any-modality human motion synthesis","AnyMo pairs FSQ tokenizer with scalable masked modeling on OmniHuMo","OmniHuMo dataset trains AnyMo for flexible multimodal motion control"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The OmniHuMo dataset supplies precisely aligned multimodal annotations at a scale and quality sufficient for one transformer to generalize across arbitrary modality combinations without modality-specific fine-tuning.","fun_headline_variants_meta":{"raw":{"variants":["AnyMo transformer generates motion from any mix of text speech and paths","Masked modeling transformer enables any-modality human motion synthesis","AnyMo pairs FSQ tokenizer with scalable masked modeling on OmniHuMo","OmniHuMo dataset trains AnyMo for flexible multimodal motion control"]},"model":"grok-4.3","cost_usd":0.002632,"raw_usage":{"total_tokens":1454,"prompt_tokens":587,"num_sources_used":0,"completion_tokens":71,"cost_in_usd_ticks":26324500,"prompt_tokens_details":{"text_tokens":587,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":796,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":587,"tokens_out":71,"duration_ms":7258,"temperature":1.0,"reasoning_tokens":796,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T08:50:29.029945+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled test in which the model receives deliberately misaligned modality labels for a held-out combination and produces lower fidelity or less controllable outputs than single-modality baselines.","supporting_citations":[],"review_version":1}