{"id":"3154b75c-050e-4238-903f-7ea2ad144c4b","arxiv_id":"2507.23143","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"X-NeMo trains a 1D identity-agnostic motion descriptor end-to-end with a diffusion model, enabling zero-shot portrait animation with improved identity and expression fidelity.","lead":"X-NeMo animates a still portrait using facial movements from a video of a different person, by compressing the motion into a small code that controls a diffusion model. The authors report better identity preservation and expression transfer than existing systems, with code promised for research.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SOTA claim rests on a self-collected benchmark with no error bars and asymmetric baseline fine-tuning; Table 1 gaps may be within noise.","rationale":"The paper is a plausible engineering contribution: the 1D-latent cross-attention design is clearly motivated, the ablations support several design choices, and the authors are transparent about limitations and about baseline availability. My concern is not with the internal mechanics but with the external validity of the SOTA claim. Table 1 is the only quantitative evidence for the headline claim, and it has three concrete weaknesses: a self-assembled test set without documented splits, no error bars or significance testing, and asymmetric baseline treatment. These are internal evidential gaps, not disagreements with external consensus. The disentanglement mechanism is also a hypothesis rather than a proof, but the empirical SOTA assertion is more directly load-bearing: if the benchmark is not reproducible, the claim fails even if the architecture works as intended. The reader already reached CONDITIONAL based partly on these evaluation concerns, so my critique does not move the verdict; it sharpens the condition that must be met before the SOTA claim is accepted.","tokens_in":15826,"tokens_out":4143,"duration_ms":48395,"concrete_test":"Release the exact test video lists, preprocessing pipeline, and evaluation code; recompute Table 1 on a public held-out benchmark (e.g., VFHQ/HDTF test splits) with every baseline run from its official checkpoint under identical preprocessing, and report mean +/- standard deviation or bootstrap confidence intervals over at least three runs. If X-NeMo no longer tops all metrics or the gaps shrink below noise, the 'consistently outperforms' claim should be softened to 'competitive on our benchmark'.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is the empirical SOTA assertion in the abstract and Section 4.2: 'consistently outperforms all baseline methods' in Table 1. The only quantitative support for this claim is a self-collected benchmark: 100 in-the-wild portraits and 300 videos from undisclosed splits, with no error bars, confidence intervals, or significance tests. The comparison is also asymmetric: AniPortrait, X-Portrait, and PD-FGC are fine-tuned on the authors' training data, while FYE, LivePortrait, and EchoMimic are used as released pretrained models. Several reported gaps are tiny (e.g., L1 0.057 vs 0.055; SSIM 0.812 vs 0.826) and could easily be run-to-run or video-set noise. The 1D-bottleneck disentanglement argument is a design hypothesis, not a proven guarantee, but even if partial identity leakage existed, the SOTA claim could still be true if the evaluation were solid. Therefore the load-bearing assumption is the validity of the benchmark and comparison protocol, not the abstraction guarantee alone.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes X-NeMo, a zero-shot diffusion-based portrait animation method. It introduces a 1D latent motion descriptor extracted from driving frames by a motion encoder and injected into a Stable-Diffusion U-Net via newly inserted cross-attention layers, alongside a reference network for appearance conditioning and temporal modules for video consistency. Training is self-supervised on talking-head and expression video datasets, with spatial/color augmentations, a dual GAN decoder head for image-level latent supervision, and reference-feature masking to reduce motion leakage from the appearance branch. The paper reports quantitative comparisons against several baselines in Table 1 on a self-collected benchmark, ablations in Table 2, and demonstrations of latent motion interpolation and portrait video outpainting.","tokens_in":16084,"tokens_out":5995,"duration_ms":67268,"significance":"If the empirical results are reliable, the paper makes a solid contribution to portrait animation: it shows that a compact 1D motion embedding with cross-attention control can drive a diffusion portrait animator, and the ablations give plausible evidence for each design choice. Strengths include end-to-end training without pretrained motion detectors, external evaluation metrics (ArcFace, MediaPipe, EmoNet) that avoid direct circularity, qualitative and quantitative ablations, and a stated commitment to release code and models. The claimed state-of-the-art status, however, rests on a self-collected benchmark whose comparison protocol is asymmetric and lacks any statistical characterization; until that evidence is strengthened, the abstract's SOTA claim is not fully supported.","major_comments":[{"comment":"The central claim that X-NeMo 'consistently outperforms all baseline methods' is supported only by point estimates on a self-collected benchmark (100 reference portraits plus 300 driving videos), with no error bars, confidence intervals, or significance tests. Several margins are small (e.g., self-reenactment L1 0.057 vs. 0.055 for AniPortrait; SSIM 0.812 vs. 0.826), and even the larger margins (e.g., EMO-SIM 0.65 vs. 0.52) could be sensitive to the particular video set and generation seed. Please report standard deviations or bootstrap confidence intervals over videos and repeated inference, run paired significance tests (e.g., Wilcoxon signed-rank) for each metric, and state exactly how many reference–driving pairs are scored in each row.","section":"§4.2, Table 1"},{"comment":"The baseline comparison is asymmetric: AniPortrait, X-Portrait, and PD-FGC are fine-tuned on the authors' training data, while FYE, LivePortrait, and EchoMimic are used as released pretrained models. With no fine-tuning budget or protocol stated, this asymmetry can change rankings, especially where Table 1 margins are within a few hundredths. Please either fine-tune all baselines under a comparable protocol, report both zero-shot and fine-tuned results for each baseline, or explicitly justify why the chosen protocol is fair; at minimum, disclose the fine-tuning data and iteration counts.","section":"§4.2, Evaluation protocol"},{"comment":"The text lists EchoMimic among the compared baselines, but Table 1 contains no EchoMimic row, so the phrase 'all baseline methods' is not supported by the table as printed. Also, the self-reenactment columns L1/SSIM/LPIPS are pixel- and feature-space image similarities, not motion-accuracy metrics, despite the text saying they assess 'image quality and motion accuracy.' Please add the missing row (or remove EchoMimic from the list) and include a direct motion metric (e.g., AED/APD or landmark distance) for self-reenactment.","section":"§4.2, Table 1 row coverage"},{"comment":"The 1D bottleneck is described as a low-pass filter that guarantees identity-motion disentanglement, but this is a design hypothesis rather than a proven property. The paper's own high-masking-ratio experiment (Appendix B) shows that when reference features are heavily masked, the motion encoder 'compensates by encoding appearance information,' indicating the bottleneck does not strictly prevent identity encoding. I recommend a direct identity-leakage probe (e.g., train a linear classifier on the motion latent to predict the driving identity, or report a mutual-information proxy) and, absent that, softening the guarantee language in the abstract and Section 3.2.","section":"§3.2, Appendix B"}],"minor_comments":[{"comment":"The sentence 'we aim to advent the field of zero-shot portrait reenactment' appears to use 'advent' where 'advance' is intended.","section":"§1"},{"comment":"There are typos: 'sorely with the diffusion loss' should be 'solely with the diffusion loss', and 'yeilding' should be 'yielding'.","section":"§4.3"},{"comment":"Please state whether the reported batch sizes are per-GPU or total, and clarify that the evaluation videos/portraits are disjoint from the training data (HDTF, VFHQ, NerSemble).","section":"§4.1"},{"comment":"The CFG negative prompt uses a motion latent extracted from the reference image; since the reference image has its own expression, this choice may suppress motion transfer when reference and driving expressions are correlated. Please justify this design or note the limitation.","section":"Eq. (2)"},{"comment":"The caption says the motion embedding is encoded from the driving image 'after applying spatial and color augmentations'; those augmentations are training-only, and the caption should make that explicit to avoid implying augmentations are used at inference.","section":"Figure 2 caption"},{"comment":"For the 'w/o cross-attn' ablation, please describe how the 1D latent is expanded into a 2D control map and which UNet layers receive the additive control, so the comparison is reproducible.","section":"Table 2, w/o cross-attn"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the closest prior baseline, X-Portrait, shares several authors with this submission and is one of the fine-tuned baselines. This makes the asymmetry of the comparison protocol and the absence of statistical tests more consequential; I would ask the authors to be especially transparent about X-Portrait's fine-tuning configuration. The architectural proposal is interesting and within scope, and the paper should be revisable without changing its core method."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know two things about this paper. First, the core design—compressing driving motion into a 1D latent vector and injecting it via cross-attention, with a dual GAN head and augmentations to push identity-motion disentanglement—is genuinely well put together. Second, the quantitative \"consistently outperforms all baselines\" claim in Table 1 is the soft spot; the benchmark is self-collected, some baselines are fine-tuned while others are not, and there are no error bars, so a few of the reported gaps look like they could be noise.\n\nWhat's new: prior work used 1D latent pose descriptors in GANs (Burkov et al., Wang et al.) and diffusion animation used spatial control or 2D attention maps (X-Portrait). Putting a 1D descriptor in an end-to-end diffusion framework with cross-attention control is a meaningful step, and the ablations back the design choices: removing the GAN head, augmentations, cross-attention, or reference feature masking all degrade results in the expected direction. The failure cases and limitations discussion are honest. The applications (interpolation, outpainting) are a bonus, not the main event.\n\nSoft spots, in order of importance. (1) The benchmark. 100 reference portraits and 300 videos, but no description of how the 300 driving videos were split or licensed beyond \"licensed videos\", no error bars, and no significance tests. Table 1 has gaps like L1 0.057 vs 0.055 and SSIM 0.812 vs 0.826; those are plausibly run-to-run noise. Also the asymmetry: AniPortrait, X-Portrait, and PD-FGC get fine-tuned on the authors' data while FYE, LivePortrait, and EchoMimic run as released pretrained models. That is not a level playing field for a state-of-the-art claim. (2) The disentanglement guarantee is a design hypothesis. A 1D bottleneck plus augmentations makes identity leakage harder but does not guarantee it; the paper implicitly acknowledges this in the masking-ratio ablation, where too-high masking causes the motion encoder to compensate by encoding appearance. So the architecture reduces leakage, it does not eliminate it. (3) Minor: the paper is already published as ICLR 2025, so the arXiv version's novelty framing reads a bit oddly, but that does not affect the technical content.\n\nOverall: the central idea is plausible and the ablations are informative. The paper deserves serious peer review; it just needs a more rigorous evaluation to support the SOTA claim. I would accept it for review with the expectation that the authors tighten the benchmark, add error bars or significance tests, and soften \"consistently outperforms\" until the evidence is stronger. This is a paper for researchers working on portrait animation or controllable diffusion; they will get useful ideas from the cross-attention control and the training strategies. I would cite it if I were working in this area.","headline":"Well-engineered portrait animation paper with a plausible architecture, but the headline SOTA claim rides on a self-collected benchmark with no error bars and uneven baseline tuning.","tokens_in":16621,"tokens_out":2305,"would_cite":true,"duration_ms":24077,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"X-NeMo claims that a compact 1D motion vector, injected through cross-attention, animates portraits without leaking the driver's identity.","keywords":["portrait animation","diffusion model","identity disentanglement","motion reenactment","cross-attention","latent motion descriptor","zero-shot","facial expression transfer"],"falsifier":"Train a linear or shallow classifier on extracted motion descriptors f_mot from many identities and test whether the driver's identity can be predicted above chance; if accuracy is well above chance, the 1D code carries identity information, contradicting the claimed identity-agnostic property. A complementary check would measure the identity similarity between generated frames and the driving identity in cross-reenactment on a large benchmark; values rivaling the reference identity similarity would indicate leakage through the motion path.","tokens_in":15624,"feed_emoji":"🎭","tokens_out":4605,"duration_ms":49153,"temperature":0.7,"pith_summary":"This paper proposes X-NeMo, a zero-shot diffusion-based portrait animation method that transfers facial motion from a driving video to a static portrait of a different person. Its central claim is that identity leakage and loss of subtle or extreme expressions come from two design choices in prior work: motion descriptors that carry spatial structure, and spatially aligned additive guidance (ControlNet-style) into the diffusion backbone. X-NeMo instead distills motion into a compact 1D latent vector, injects it through cross-attention, and trains end-to-end with a dual GAN decoder plus color and spatial augmentations. The authors report that this outperforms current baselines on self- and cross-reenactment, with the best identity similarity and emotion similarity scores on their benchmark.","feed_headline":"One compact motion vector animates faces without identity leakage","feed_subtitle":"X-NeMo injects motion via cross-attention and reports top identity and emotion scores across self- and cross-reenactment benchmarks.","key_machinery":"The load-bearing object is the implicit 1D latent motion descriptor f_mot, produced by a motion encoder E_mot (a feature-alignment backbone with attention layers and MLP heads) from the driving image. Its compactness is meant to act as a low-pass information bottleneck that excludes 2D structural cues; motion is injected into the diffusion U-Net through newly inserted cross-attention layers, so no spatially aligned additive offset reaches the backbone. A dual GAN decoder head (a StyleGAN-style generator) co-trained with image-level losses guides the descriptor toward fine-grained expressions, while spatial and color augmentations and 30% reference-feature masking push identity and motion apart. A relative translation and scale triplet (Δx, Δy, scale ratio) accounts for head motion lost by face-centered cropping.","core_discovery":"X-NeMo's central claim is that a structure-agnostic motion control path, built on a 1D identity-agnostic latent motion descriptor, can drive a pretrained latent diffusion model to perform expressive zero-shot portrait animation while preserving the reference identity. The authors argue that explicit motion signals such as landmarks and synthetic cross-identity images encode the driving identity's structure, and that ControlNet-like additive spatial guidance lets the U-Net shortcut semantic correspondence by mimicking 2D layout, both causing identity leakage. Their remedy is an end-to-end motion encoder that outputs a 512-dimensional global latent, cross-attention injection into the U-Net, color and spatial augmentation of driving frames, reference-feature masking, and a jointly trained GAN head that supervises the latent with image-level losses. On their benchmark, the method reports the best L1, SSIM, LPIPS, ID-SIM, AED/APD, and EMO-SIM among the methods compared.","pith_inferences":["If the 1D bottleneck truly blocks identity, the same structure-agnostic cross-attention control could generalize to full-body or object animation, where 2D pose conditions typically leak source identity or structure.","The identity-agnostic claim is directly testable: a probe classifier trained on extracted motion descriptors should not predict the driver's identity above chance level.","The paper's framing implies that residual identity leakage in any diffusion animator can be diagnosed by asking whether its motion control path carries spatial structure, and that such leakage may be reduced by compressing the motion condition into a global latent."],"forward_implications":["Cross-identity reenactment should preserve the reference identity even when the driving and reference faces differ strongly in structure, style, and appearance.","The motion descriptor supports latent motion interpolation and video outpainting, so it can serve as a unified motion representation beyond frame-to-frame animation.","End-to-end training without pretrained motion detectors means the system can improve as more diverse and expressive video data become available.","Classifier-free guidance that uses the reference's own motion as a negative prompt steers inference toward more accurate expression transfer.","Replacing spatially aligned additive control with cross-attention to a global latent may be a general recipe for reducing conditional leakage in diffusion models."],"supporting_citations":[{"why":"Supplies the latent pose descriptor and low-pass bottleneck idea, plus the GAN-head loss formulation.","marker":"Burkov et al. (2020)"},{"why":"Provides the 1D latent motion representation and disentangled training scheme that X-NeMo adapts.","marker":"Wang et al. (2022; 2023)"},{"why":"Motivates the information bottleneck principle behind using a smaller motion encoder and lower-dimensional latent.","marker":"Tishby et al. (2000)"},{"why":"Supplies the reference network with mutual self-attention for appearance conditioning.","marker":"Cao et al. (2023)"},{"why":"Provides the latent diffusion model backbone that X-NeMo builds on.","marker":"Rombach et al. (2022)"},{"why":"Defines the ControlNet spatially additive control scheme that X-NeMo contrasts with and ablates against.","marker":"Zhang et al. (2023b)"},{"why":"X-Portrait is a key baseline whose synthetic cross-identity training pairs and motion attention are compared and critiqued.","marker":"Xie et al. (2024)"},{"why":"Inspires the reference feature masking strategy for reducing motion leakage in appearance features.","marker":"He et al. (2022)"},{"why":"Supplies the StyleGAN generator architecture used as the dual GAN head.","marker":"Karras et al. (2020)"},{"why":"Provides the classifier-free guidance formula used at inference with the reference's motion as negative prompt.","marker":"Ho & Salimans (2022)"}],"fun_headline_variants":["1D motion vector animates faces, preserving identity","Cross-attention motion injection stops identity leak in face reenactment","One compact latent captures face motion, no identity leakage","Expressive face reenactment from a single 512-D motion vector","Zero-shot face animation with disentangled motion and identity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method's disentanglement rests on the assumption that a 1D bottlenecked latent vector, even with augmentations and a GAN decoder, cannot encode identity-specific spatial structure; if the motion encoder packs identity information into that vector, identity leakage will persist despite the architecture.","fun_headline_variants_meta":{"raw":{"variants":["1D motion vector animates faces, preserving identity","Cross-attention motion injection stops identity leak in face reenactment","One compact latent captures face motion, no identity leakage","Expressive face reenactment from a single 512-D motion vector","Zero-shot face animation with disentangled motion and identity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000619,"raw_usage":{"total_tokens":2887,"prompt_tokens":973,"completion_tokens":1914,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":589,"completion_tokens_details":{"reasoning_tokens":1829}},"tokens_in":589,"tokens_out":1914,"duration_ms":14901,"temperature":1.0,"reasoning_tokens":1829,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:59:54.614575+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a linear or shallow classifier on extracted motion descriptors f_mot from many identities and test whether the driver's identity can be predicted above chance; if accuracy is well above chance, the 1D code carries identity information, contradicting the claimed identity-agnostic property. A complementary check would measure the identity similarity between generated frames and the driving identity in cross-reenactment on a large benchmark; values rivaling the reference identity similarity would indicate leakage through the motion path.","supporting_citations":[],"review_version":1}