{"id":"54411288-938e-4f34-b1bc-ea5fbf2cc71b","arxiv_id":"2607.10313","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Dyadic head motion is generated by modulating frozen monologic speech-to-motion priors with interaction latents predicted from paired audio via monologic-anchored factorization and cross-attentive dual-stream encoding.","lead":"Learn2Chat builds two-person talking-head motion from audio by first running each speaker through a frozen single-speaker face model, then applying a learned interaction tweak so the two faces react to each other. That reuse of existing speech-to-face models makes interactive avatars cheaper to train and more coherent when conversation data is scarce.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The monologic-neutral residual may not be pure interaction; DualTalk monologic bases already carry conversational style, so SOTA gains may partly be residual fitting rather than clean modulation.","rationale":"The reader correctly flags the monologic-neutral residual as the weakest assumption and already rates the paper CONDITIONAL for fairness/single-benchmark reasons. The stress test sharpens the same point with a concrete DualTalk-specific failure mode: monologic bases are not interaction-neutral because DualTalk audio is conversational, so Lswap can encode non-interaction residuals. No internal contradiction or math error is present; tables, ablations, user study, and cross-backbone supplement still support a useful modular system. The concern does not justify REJECT, but it keeps the verdict CONDITIONAL until a content-vs-interaction probe (or true monologic-only pretraining of the factorization) is shown. Agreement with the reader is therefore agree; no verdict change beyond confirming CONDITIONAL.","tokens_in":22455,"tokens_out":684,"duration_ms":7185,"concrete_test":"On held-out DualTalk clips, compute MI or linear predictability of zd (from EI) from self-speech Wav2Vec2 features alone (vs. partner features). Then content-swap: replace speaker A’s audio content while keeping partner B fixed, re-encode z, and measure whether generated motion of A changes mainly in interaction cues (rPCC, listener FD) or also in lip/jaw MSE. If self-speech alone predicts zd strongly, or content-swap moves speech-aligned MSE more than coordination metrics, the residual is not pure interaction and the load-bearing assumption fails.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that monologic motion is an interaction-neutral baseline and that the residual after Monologic-Anchored Factorization (Sec. III-B: shared ES, learnable zm ~ N(μm,σm²), Lswap, Lsem-cyc/Lint-cyc) is a clean, audio-predictable interaction latent z that AdaLN can inject without relearning speech-motion mapping. DualTalk is multi-round conversational data; monologic bases are obtained by running Gmono independently on each DualTalk audio stream (Eq. 1). Those streams are not monologic speech: they contain turn-taking, backchannels, and partner-conditioned prosody. Consequently xm already embeds interaction-correlated dynamics. Lswap then forces D(sm,zd)≈xd, so zd can absorb any residual that makes monologic-looking motion look dyadic, including speech-content residuals that monologic Gmono failed to capture on conversational audio, not only partner-responsive social feedback. The paper never shows that z is independent of speech content (no mutual-information or content-swap probe), and the strongest ablations (Table IV w/o factorization; Table V loss drops) only show that the factorization helps reconstruction/metrics, not that the residual is pure interaction. If the residual is contaminated, the “interaction modulation over frozen monologic priors” story and the claimed data-efficient transfer across backbones are weaker than stated; gains could be residual fitting on DualTalk rather than structured separation.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"Learn2Chat reformulates audio-driven dyadic 3D head-motion generation as interaction modulation over frozen pretrained monologic generators. A Monologic-Anchored Motion Factorization (shared semantics encoder ES, stochastic interaction encoder EI for dyadic motion, learnable monologic interaction prior zm, and AdaLN decoder) is trained with reconstruction, swap, cycle, and KL losses to obtain an interaction latent z. A dual-stream cross-attentive audio encoder then predicts za from paired Wav2Vec2 features, aligned to motion-derived zd via InfoNCE, and modulates monologic semantics at inference. On DualTalk (test and OOD), with DEEPTalk and UniTalker backbones, the method reports SOTA FD/P-FD/MSE/rPCC, favorable parameter count, a user study (n=25), ablations, listener-role analysis, and cross-backbone transfer.","tokens_in":22965,"tokens_out":958,"duration_ms":24986,"significance":"If the empirical gains hold under fair audio-only deployment, the paper offers a practical and timely alternative to large jointly trained dyadic generators: reuse strong monologic speech–motion models and learn a comparatively lightweight interaction adapter. The dual-stream cross-attention design, two-stage freeze-then-predict protocol, DualTalk SOTA tables, user study, parameter comparison, and backbone-swap results in the supplement are concrete strengths for TVCG-style conversational avatar work. Even if the residual is not a pure social-interaction factor, a residual-modulation recipe that improves coordination metrics while remaining plug-and-play would still be useful to the community.","major_comments":[{"comment":"Sec. III-A–B, Eqs. (1)–(10): the central narrative treats xm = Gmono(Ai) as an interaction-neutral canonical baseline and zd as clean partner-induced modulation. DualTalk streams are multi-round conversational audio (turn-taking, backchannels, partner-conditioned prosody), so Gmono(Ai) already sees interaction-correlated speech; Lswap then forces D(ES(xm), zd) ≈ xd, so zd can absorb monologic prediction error and speech-content residuals as well as social feedback. The manuscript never tests speech–interaction independence (e.g., content-swap, mutual information between z and phoneme/text features, or holding A fixed while swapping partner context). Please either add such probes or substantially soften claims of “clean interaction representations” / “separates intrinsic speech-driven motion from social interaction effects” throughout abstract, intro, and conclusion.","section":null},{"comment":"Sec. III-B factorization data construction is under-specified relative to the monologic-anchor claim. It is unclear whether xm used in Lrec/Lswap is (i) Gmono applied to DualTalk conversational audio only, (ii) real monologic corpus motion, or (iii) DualTalk segments treated as monologic. If (i), the “semantic manifold learned from monologic data” is largely the monologic model’s residual on dyadic speech, which weakens the prior-reuse story and the interpretation of Table V. Clarify the exact sources of xm/xd, whether Gmono is frozen DualTalk-retrained or original monologic pretraining, and report a control that freezes ES/D trained only on true monologic data if available.","section":null},{"comment":"Sec. IV-B baseline protocol: L2L, DIM, and DualTalk are evaluated audio-only by substituting partner GT motion with monologic base motions. That adaptation is deployment-relevant but systematically handicaps methods designed for GT partner motion; UniLS is the only native audio-only dyadic peer. Please report (or clearly mark as oracle) the original motion-conditioned DualTalk/L2L/DIM numbers with GT partner motion alongside the adapted setting, and avoid framing large gaps vs adapted DualTalk as pure evidence that joint dyadic learning is inferior. The UniLS comparison and parameter table (Table III) should carry more of the SOTA argument.","section":null},{"comment":"Table IV OOD row for Ours lists jaw MSE as 0.60 while Tables I–II and the same table’s other rows place jaw MSE near 1.4; this looks like a transcription error and affects ablation interpretation. Please audit all metric tables (including supplement listener/cross-backbone tables) for consistency of scaling (×10−k) and values, and re-state which checkpoint/split each ablation uses.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: freeze a good monologic speech-to-motion model, factor dyadic motion into shared semantics plus a stochastic interaction latent, then predict that latent from paired audio with dual-stream cross-attention and AdaLN modulation. On DualTalk that beats monologic baselines, adapted motion-conditioned methods, and UniLS on FD/P-FD/MSE/rPCC (test and OOD) with DEEPTalk and UniTalker, plus a cleaner user-study win on interaction naturalness. Parameter count is modest relative to DualTalk/UniLS, and the supplement shows cross-backbone swap and better listener-side numbers.\n\nWhat is actually new is the engineering of the separation, not a new generative family. Shared ES on monologic and dyadic motion, learnable monologic “neutral” prior zm, swap + cycle losses, then frozen factorization while training only the audio-to-z path with recon + InfoNCE. That is a clearer prior-reuse story than UniLS’s joint latent or DualTalk’s partner-motion conditioning. Ablations (Tables IV–V) move in the right direction when factorization or dual-stream CA is removed, and the math is ordinary VAE/Transformer/AdaLN with no internal contradiction.\n\nThe stress-test lands partially. DualTalk audio is conversational, so Gmono(A) already carries turn-taking and partner-conditioned prosody; Lswap can let z absorb any residual that turns monologic-looking motion into dyadic motion, not only pure social feedback. There is no MI or content-swap probe that z is speech-independent. So the “clean interaction modulation” narrative is stronger than the evidence; gains could partly be residual fitting on this benchmark. Fairness of the adapted L2L/DIM/DualTalk baselines is also a bit soft (monologic motion as partner condition is not their native setting). Single-benchmark, head-only, no code/error bars—standard for the area, not fatal.\n\nThis is for people building conversational avatars who already have monologic backbones and want a lightweight interaction adapter. The central empirical claim holds; the interpretation is a bit oversold. I would send it to referees and would cite the DualTalk numbers and the modular recipe if I were shipping a similar system. Worth a reading-group slot if the group cares about talking heads.","headline":"Solid modular reformulation of dyadic head motion as modulation over frozen monologic priors; DualTalk SOTA is real, but the “clean interaction residual” story is only partly proven.","tokens_in":23546,"tokens_out":608,"would_cite":true,"duration_ms":5969,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Dyadic talking-head motion is interaction modulation of frozen monologic speech-motion priors, not a new joint generator.","keywords":["dyadic motion generation","audio-driven facial animation","monologic priors","interaction modulation","motion factorization","talking heads","digital humans"],"falsifier":"If swapping the monologic backbone at test time (or ablating the factorization stage) collapses inter-speaker correlation and Fréchet distances back to monologic or non-factorized levels on DualTalk, the claim that the residual is a transferable interaction latent fails.","tokens_in":23362,"feed_emoji":"🗣️","tokens_out":640,"duration_ms":6596,"temperature":0.7,"pith_summary":"Most systems that animate two people in conversation either need the partner's ground-truth motion or train a single model that mixes speech-driven facial motion with social feedback. This paper argues that the speech-to-motion map is already well learned by monologic talking-head models, so the real missing piece is only the residual social adjustment. Learn2Chat freezes a pretrained monologic generator, factors dyadic sequences into speech semantics plus an interaction latent, and then predicts that latent from the two audio streams alone. The predicted latent modulates the monologic base motion through the decoder, producing coordinated head motion without relearning lip sync from scratch. On the DualTalk benchmark the method leads quantitative metrics and user preference, and the same interaction module works when the monologic backbone is swapped at test time.","feed_headline":"Two-person head motion is just monologic motion plus a social residual","feed_subtitle":"Freeze a single-speaker talking-head model, predict the partner residual from dual audio, and get coordinated conversation.","key_machinery":"Monologic-Anchored Motion Factorization: a shared semantics encoder plus an asymmetric interaction encoder that maps monologic motion to a learnable neutral prior and dyadic motion to a stochastic residual; swap and cycle losses force the residual to carry only the domain gap. That residual is later predicted from dual audio by cross-attentive streams and injected through AdaLN into the frozen monologic decoder.","core_discovery":"Dyadic conversational head motion is better generated by modulating a frozen, pretrained monologic speech-motion prior with an audio-predicted interaction latent than by training a unified dyadic generator that entangles self-speech and partner feedback. The Monologic-Anchored Motion Factorization isolates a clean interaction residual; Cross-Attentive Interaction Latent Prediction recovers that residual from paired speech; the residual then steers the monologic base via AdaLN, yielding higher fidelity, better inter-speaker coordination, and backbone-agnostic reuse.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Dyadic heads: monologic motion plus dual-audio social residual","Freeze monologic talking-head prior, modulate with interaction latent","Paired speech predicts residual that steers monologic motion into chat","Interaction latents over monologic priors beat unified dyadic generators","Monologic-anchored factorization yields clean residual for two-person heads"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That monologic head motion is an interaction-neutral baseline whose residual after factorization is a pure, audio-predictable social modulation that the monologic manifold can cleanly separate from speech content.","fun_headline_variants_meta":{"raw":{"variants":["Dyadic heads: monologic motion plus dual-audio social residual","Freeze monologic talking-head prior, modulate with interaction latent","Paired speech predicts residual that steers monologic motion into chat","Interaction latents over monologic priors beat unified dyadic generators","Monologic-anchored factorization yields clean residual for two-person heads"]},"model":"grok-4.5","effort":"low","cost_usd":0.00394,"raw_usage":{"total_tokens":1264,"prompt_tokens":863,"num_sources_used":0,"completion_tokens":73,"cost_in_usd_ticks":39400000,"prompt_tokens_details":{"text_tokens":863,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":328,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":863,"tokens_out":73,"duration_ms":4087,"temperature":1.0,"reasoning_tokens":328,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T12:41:51.217905+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If swapping the monologic backbone at test time (or ablating the factorization stage) collapses inter-speaker correlation and Fréchet distances back to monologic or non-factorized levels on DualTalk, the claim that the residual is a transferable interaction latent fails.","supporting_citations":[],"review_version":1}