{"id":"fce671b8-4137-416b-b55d-62e3862b1f85","arxiv_id":"2505.09827","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Dyadic Mamba uses a state-space backbone and simple concatenation to generate text-conditioned two-person motion that extends beyond training length, with a per-person NDMS long-term benchmark.","lead":"This paper presents a Mamba-based diffusion model that generates two-person motion from text and keeps generating beyond the 10-second clips it was trained on. The authors also propose a long-term evaluation benchmark, but it measures only single-person motion quality, not the quality of the interaction itself.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Long-term dyadic claim leans entirely on per-person local NDMS; interaction coherence is unmeasured, so the claimed superiority over transformers for long-term dyadic motion is not quantitatively established.","rationale":"Reading in good faith, the paper makes a plausible, simple architectural contribution: replacing transformer backbones with Mamba and exchanging cross-attention for concatenation achieves competitive short-term results on InterHuman and Inter-X, with fewer parameters, and produces visually stable long generations. The central claim, however, is specifically about long-term dyadic synthesis that extrapolates beyond the 10s training horizon. The load-bearing evidence for that claim is Table 2, and the authors themselves restrict it to per-person individual motion quality. A metric computed on 1/3s windows per person cannot verify interaction coherence, which is the defining property of dyadic motion. The reader's weakest-assumption analysis identifies exactly this gap; I agree. The concern is not that the architecture is internally inconsistent, but that the evaluation does not support the strongest reading of the claim. The paper's own limitation statement in Section 4.4 is the clearest evidence of the mismatch, and Figure 7 confirms that interaction-level failures occur. The Table 1 Inter-X Top-1 value of 3.658 is additionally an impossible R-Precision score (accuracy > 1) and weakens confidence in the reported numbers, though it does not by itself overturn the short-term results. A concrete interaction-focused evaluation would settle whether the long-term dyadic advantage is real. Since the reader already recommended CONDITIONAL and the concern does not contradict that verdict, no change is needed.","tokens_in":12684,"tokens_out":4062,"duration_ms":45998,"concrete_test":"Generate 28s sequences from Dyadic Mamba, InterGen, and InterGen (RoPE) on a fixed set of 100 InterHuman prompts with the released models. Compute interaction-aware metrics per frame and per 1s sliding window: minimum distance between the two SMPL body meshes (contact maintenance and penetration rate), relative facing angle between persons, inter-person distance distribution versus real 10s samples, and identity/order consistency across windows. In parallel, run a small human study on short video clips asking annotators to rate contact plausibility and interaction coherence without knowing the method. If Dyadic Mamba does not significantly beat InterGen on these metrics, the claim should be narrowed to single-person long-term motion quality. As a secondary check, repair the Table 1 Inter-X Top-1 R-Precision value (3.658 is out of range; likely 0.365) and confirm the corrected table.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.4 explicitly concedes: \"our approach only evaluates individual motion quality, and per-frame human interaction evaluation remains an open problem for future work.\" Table 2 and Figure 4 nonetheless carry the central long-term claim, but the reported metric — per-person per-frame NDMS over 1/3-second windows — is local, directional, and single-person. It cannot detect the failure modes most relevant to dyadic motion: body penetration, loss of contact, facing/role violations, person-order swaps, or long-horizon interaction structure. Figure 7 even shows such failures in the proposed model (clipping, order changes, one person walking while the other is instructed to interact). The headline statement that Dyadic Mamba \"significantly outperforms transformer-based approaches on longer sequences\" is therefore supported only by qualitative examples (Figure 1) for the specifically dyadic aspect of the claim. What Table 2 actually establishes is that local single-person motion quality degrades less for Dyadic Mamba than for InterGen beyond 10s. That is a meaningful result, but it is not the claimed result. The absence of a real-data reference at 14s and 28s in Table 2 further means the metric only measures closeness to short-window training-motion statistics, not long-sequence realism or interaction coherence. The paper's own caveat thus undercuts the strongest interpretation of its central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Dyadic Mamba, a diffusion-based state-space model for text-to-dyadic human motion synthesis. The architecture processes the two persons' motion sequences with shared per-person Mamba blocks and exchanges information through concatenation with a learnable down-projection, avoiding cross-attention. The paper reports competitive short-term results on InterHuman and Inter-X, and proposes a long-term benchmark based on per-person per-frame Normalized Directional Motion Similarity (NDMS). The central claim is that Dyadic Mamba extrapolates beyond the 10-second training horizon and 'significantly outperforms transformer-based approaches' on longer dyadic sequences, while remaining competitive on short-term benchmarks. The paper also includes ablations of model size, conditioning, and cross-person information flow, and discusses failure cases.","tokens_in":12987,"tokens_out":3600,"duration_ms":36530,"significance":"The architectural direction is timely: demonstrating that a linear-time SSM backbone can handle dyadic motion synthesis and extrapolate temporally is a useful contribution, especially if it holds with parameter efficiency. The authors are also transparent in reporting a long-term evaluation protocol, which is valuable for future work. However, the strongest advertised claim — significant superiority over transformers for long-term dyadic synthesis — is not quantitatively established: the long-term metric measures only individual per-person motion quality, and the paper explicitly concedes that interaction quality is not evaluated. The short-term Inter-X results also contain an impossible reported number. If the metric issues and reporting errors are corrected and the claims are re-scoped to local single-person motion quality over long horizons, the contribution is solid but less headline-worthy. The use of NDMS from the first author's prior work is not circular, since it is an external quality metric rather than a learned evaluator fitted to the model, but it is not sufficient for the dyadic claim.","major_comments":[{"comment":"The reported Inter-X Top-1 R-Precision for the proposed method is 3.658±0.007, which is larger than 1 and therefore cannot be an accuracy. This is an impossible value for a retrieval-based precision metric and makes the Inter-X short-term comparison uninterpretable as printed. The authors must correct the value (likely a decimal error, e.g., 0.3658) and re-evaluate the associated conclusions about competitive performance on Inter-X.","section":"Table 1, Inter-X block"},{"comment":"The central long-term claim is not supported by the reported metric. NDMS is computed per person over 1/3-second windows and measures local single-person motion similarity, not dyadic interaction quality. The paper itself states in Section 4.4: 'our approach only evaluates individual motion quality, and per-frame human interaction evaluation remains an open problem for future work.' The failure cases in Figure 7 (clipping, person-order changes, one person walking while the other is instructed to interact) are exactly the kinds of dyadic coherence failures that NDMS cannot detect. Consequently, Table 2 establishes only that per-person local motion quality degrades less for Dyadic Mamba than for InterGen beyond 10s; it does not establish the abstract's claim of 'significantly outperforming transformer-based approaches on longer sequences' for dyadic motion synthesis.","section":"Section 4.4, Table 2, Figure 4"},{"comment":"The long-term comparison lacks both a real-data reference at 14s and 28s and any significance testing. The 'Real' row is given only at 7s (0.451±0.152), and at 14s and 28s the models are compared only against each other. Given the reported standard deviations — for example Ours 0.379±0.157 versus InterGen 0.290±0.118 at 14s, and Ours 0.376±0.155 versus InterGen (RoPE) 0.212±0.009 at 28s — the differences are not obviously significant, and the phrase 'significantly outperforms' is unsupported. The authors should report significance tests or confidence intervals, and ideally provide real-data NDMS values at the longer horizons to calibrate what good long-term performance means.","section":"Table 2, long-term benchmark"}],"minor_comments":[{"comment":"Reference [13] contains a typo: 'Internantional Conference' should be 'International Conference on Learning Representations.'","section":"References"},{"comment":"Several citation labels in Table 1 appear inconsistent with the reference list: 'ComMDM [32]' and 'RIG [32]' both point to reference [32], but the reference list assigns [32] to the PriorMDM paper, while RIG appears to be reference [36]. Please verify all citation keys in the table.","section":"Table 1"},{"comment":"The caption says 'single-step denoising' while the model is trained with T=1000 diffusion steps and sampled with 50 DDIM steps; consider rewording to 'one denoising network evaluation' to avoid confusion.","section":"Figure 2"},{"comment":"The column header row for the ablation table is difficult to parse, especially the '#Param S M L' and 'Prepending+ ⊕' entries; a clearer layout with explicit sub-headers would improve readability.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has a solid architectural contribution, but the headline long-term dyadic claim is currently over-stated relative to the evidence. The impossible R-Precision value in Table 1 is a simple reporting error, but it must be fixed before any further consideration. I would also gently note that the long-term evaluation metric comes from the first author's prior work; while this is not circular, the paper would benefit from an independent interaction-aware metric or at least a clear statement that the proposed benchmark covers only individual motion quality. The paper fits the scope of the journal, but the revision needs to either add interaction-level evaluation or explicitly re-scope the central claim to local per-person motion quality over long horizons."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a clean, useful extension of Mamba to text-to-dyadic motion, with a sensible concatenation design and a new long-term benchmark. But the headline claim—arbitrary-length dyadic motion that beats transformers—isn't supported by the evidence they actually report. The long-term metric is per-person local NDMS, not interaction quality, and the authors concede that in Section 4.4. Figure 7 shows the model failing on the very interaction-level failure modes that the metric can't see.\n\nWhat's genuinely new: using SSM Mamba for dyadic motion with simple concatenation for cross-person communication, avoiding cross-attention, with shared weights and AdaLN conditioning. The ablation in Table 3 is careful, and the model is parameter-efficient. On InterHuman, short-term R-Precision is competitive with InterMask, and it beats data-space transformer baselines. The proposed long-term evaluation benchmark (per-person NDMS at 7s/14s/28s) is a reasonable first step, even if limited.\n\nSoft spots: the Inter-X Top-1 R-Precision of 3.658 in Table 1 is impossible—it should be between 0 and 1. That looks like a typo, but it erodes confidence in the table. More substantively, the long-term comparison has no real-data reference at 14s and 28s, and it only measures local single-person motion quality, not contact, facing, role, or interaction structure. So the paper actually shows that Dyadic Mamba's per-person motion stays smoother than InterGen's beyond 10s, not that it generates better long-term dyadic interactions. The concurrent TIMotion is cited but not compared, which is fine for a v1 but worth noting.\n\nI'd send this to a serious referee. The fix is straightforward: correct the typo, reframe the long-term claims as 'local motion quality', and ideally add at least one interaction-level metric or a qualitative user study. With that, it's a solid contribution for the text-to-motion subfield. Without the reframing, the central claim overreaches.","headline":"Useful Mamba-based recipe for dyadic motion with a clever concatenation design, but the long-term superiority claim rests on a metric that doesn't measure interaction quality and a table typo.","tokens_in":13519,"tokens_out":2914,"would_cite":true,"duration_ms":27671,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a Mamba-based diffusion model can synthesize text-described two-person motion with stable per-person quality far beyond the 10-second training horizon, where transformer-based generators degrade.","keywords":["human motion synthesis","dyadic interaction","state-space models","Mamba","text-to-motion","long-term generation","diffusion models","motion quality benchmark"],"falsifier":"Generate 28-second motions with Dyadic Mamba on InterHuman test prompts and measure per-frame contact distance between the two bodies, the angle between their facing directions, and the consistency of interaction roles; compare these against ground-truth motions. If per-person NDMS stays around 0.37 while these interaction measures deviate from the real data just as much as InterGen's do, the claim that Dyadic Mamba outperforms transformers on long-term dyadic synthesis is falsified.","tokens_in":12487,"feed_emoji":"🕺","tokens_out":6637,"duration_ms":63566,"temperature":0.7,"pith_summary":"This paper tries to establish that state-space models, not transformers, are the right backbone for generating text-described two-person (dyadic) motion, especially for long interactions. It introduces Dyadic Mamba, a diffusion model that runs each person's motion through its own Mamba layers and mixes the two streams by simple concatenation, and argues this reaches transformer-level quality on standard short-term benchmarks while keeping motion quality stable at 14 and 28 seconds—beyond the 10-second training length. If true, long social interactions could be synthesized in one pass without windowing tricks or positional-encoding fixes, and with fewer parameters. The paper also contributes a long-term evaluation benchmark based on per-person motion similarity measured over one-third-second windows.","feed_headline":"State-space model keeps two-person motion stable past training length","feed_subtitle":"Dyadic Mamba holds per-person quality at 14 and 28 seconds; transformer baselines degrade.","key_machinery":"The load-bearing object is the Mamba layer, a selective state-space model that processes a sequence in linear time by updating a recurrent hidden state, so it has no positional encoding and no quadratic attention cost; in this paper it runs as two paired modules per block—a self-Mamba per person and a cross-Mamba that receives the concatenation of both persons' intermediate features. The same weights process both persons, and Adaptive LayerNorm injects the text embedding and diffusion step. This design is trained as a data-space denoising diffusion model, and the long-term claim is measured by per-person Normalized Directional Motion Similarity over one-third-second windows at 7, 14, and 28 seconds.","core_discovery":"On its own terms, the paper's discovery is that a state-space backbone removes the length ceiling for dyadic motion generation. Dyadic Mamba is a denoising diffusion model whose denoiser is built from Mamba layers arranged in cooperative blocks: each person's motion passes first through a shared self-Mamba, the two streams are concatenated along the feature dimension and down-projected, and a second shared Mamba mixes them; conditioning on text and diffusion time step uses Adaptive LayerNorm. Trained on at most 10-second sequences, the model matches transformer-based short-term results and maintains an NDMS score around 0.37 at 7s, 14s, and 28s, whereas InterGen drops from 0.34 to 0.27 and InterGen with rotary positional embeddings drops from 0.32 to 0.21. The paper takes this as evidence that SSM-based architectures extrapolate to arbitrary length, eliminating the need for windowing or positional-encoding workarounds.","pith_inferences":["Editorial inference: the same concatenation-based interaction module may extend to three or more persons; the paper notes the addition variant is order-invariant, and an asymmetric concatenation could be replaced by a permutation-equivariant mixing operation.","Editorial inference: the benchmark's per-person focus leaves room for an interaction-level long-term metric; a fair next test would track contact distance, facing angle, and role consistency over 28-second generations.","Editorial inference: 'arbitrary length' is demonstrated only up to 28 seconds; genuinely arbitrary length would require checking whether the recurrent state saturates or drifts over minute-scale generations, which the paper does not report.","Editorial inference: the method's success suggests state-space models might also stabilize long single-person motion generation and motion in-betweening, since the length bottleneck is in the backbone, not the dyadic task."],"forward_implications":["Long social interactions can be generated in a single diffusion pass rather than stitched windows, because the state-space layer carries context without positional encodings.","Dyadic coordination does not require cross-attention: a shared Mamba over concatenated person streams is enough, which simplifies the architecture and roughly halves the parameter count relative to InterGen.","The known failure of rotary embeddings beyond roughly twice the training length transfers from language modeling to motion, explaining why the RoPE-enhanced transformer still collapses at 28 seconds.","The new per-person NDMS benchmark gives later methods a cheap, standardized way to report long-term stability on InterHuman.","Because Mamba's recurrence is linear in sequence length, generating minute-long interactions avoids the memory growth that attention would incur."],"supporting_citations":[{"why":"Provides the Mamba state-space layer that the whole generator is built from and that removes positional-encoding length limits.","marker":"[8]"},{"why":"Supplies the InterHuman dataset, the data-space diffusion baseline InterGen, and the short-term evaluation protocol the paper builds on.","marker":"[18]"},{"why":"Provides the per-person NDMS metric used to quantify long-term motion quality in the proposed benchmark.","marker":"[37]"},{"why":"Defines rotary positional embeddings, used to construct the InterGen (RoPE) baseline that tests whether better positional encoding alone fixes extrapolation.","marker":"[35]"},{"why":"Supplies the Inter-X dataset and protocol, showing the method transfers to articulated SMPL-X poses with finger motion.","marker":"[46]"},{"why":"Is the discrete state-space masked-modeling baseline compared on short-term dyadic benchmarks.","marker":"[13]"}],"fun_headline_variants":["Dyadic Mamba: SSM beats transformers at long two-person motion","State-space model extends dyadic motion synthesis to arbitrary length","Simple concatenation in SSM powers long-term dyadic motion generation","Forget cross-attention: Mamba stream concatenation makes long motion work","New benchmark shows SSM maintains quality at 28s, transformers crumble"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that per-person motion quality, averaged over one-third-second windows, is a sufficient indicator of long-term dyadic quality; if the interaction between the two people can degrade while each individual's short windows look fine, the paper's long-term advantage over transformers is not proven.","fun_headline_variants_meta":{"raw":{"variants":["Dyadic Mamba: SSM beats transformers at long two-person motion","State-space model extends dyadic motion synthesis to arbitrary length","Simple concatenation in SSM powers long-term dyadic motion generation","Forget cross-attention: Mamba stream concatenation makes long motion work","New benchmark shows SSM maintains quality at 28s, transformers crumble"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000221,"raw_usage":{"total_tokens":1435,"prompt_tokens":915,"completion_tokens":520,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":426}},"tokens_in":531,"tokens_out":520,"duration_ms":5231,"temperature":1.0,"reasoning_tokens":426,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:23:10.040099+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate 28-second motions with Dyadic Mamba on InterHuman test prompts and measure per-frame contact distance between the two bodies, the angle between their facing directions, and the consistency of interaction roles; compare these against ground-truth motions. If per-person NDMS stays around 0.37 while these interaction measures deviate from the real data just as much as InterGen's do, the claim that Dyadic Mamba outperforms transformers on long-term dyadic synthesis is falsified.","supporting_citations":[{"cited_title":"Intergen: Diffusion-based multi-human motion genera- tion under complex interactions","cited_arxiv_id":null,"evidence_quote":"Supplies the InterHuman dataset, the data-space diffusion baseline InterGen, and the short-term evaluation protocol the paper builds on."},{"cited_title":"Intention- based long-term human motion anticipation","cited_arxiv_id":null,"evidence_quote":"Provides the per-person NDMS metric used to quantify long-term motion quality in the proposed benchmark."},{"cited_title":"Roformer: Enhanced transformer with rotary position embedding","cited_arxiv_id":null,"evidence_quote":"Defines rotary positional embeddings, used to construct the InterGen (RoPE) baseline that tests whether better positional encoding alone fixes extrapolation."},{"cited_title":"Inter-x: Towards versatile human- human interaction analysis","cited_arxiv_id":null,"evidence_quote":"Supplies the Inter-X dataset and protocol, showing the method transfers to articulated SMPL-X poses with finger motion."},{"cited_title":"Intermask: 3d human interaction generation via collaborative masked modelling","cited_arxiv_id":null,"evidence_quote":"Is the discrete state-space masked-modeling baseline compared on short-term dyadic benchmarks."}],"review_version":1}