{"id":"aeebcc28-1719-4761-abb1-c260e9d9c24f","arxiv_id":"2607.24124","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Fine-tuning a video diffusion model on dense 3DMM normal maps from Pixel3DMM lets a single portrait photo be animated with a driving video's expressions and pose, surpassing landmark- and latent-based portrait animation baselines on most benchmark metrics.","lead":"ViDS is a method for animating a single face photo using another video's facial movements, powered by a video diffusion model conditioned on 3D face-track geometry. It reports better expression and pose control than prior portrait-animation methods, though some evaluation metrics overlap with the 3D conditioning signal.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Motion metrics are confounded and don't support the headline: AED/APD come from a 3DMM estimator matched to ViDS's conditioning, and on those metrics ViDS is not the best on expression distance in any table.","rationale":"The paper is well-engineered and the ablations are extensive, but the central motion-control claim is supported mainly by metrics that are computed with a 3DMM estimator (Deep3DFaceRecon) similar in kind to the Pixel3DMM conditioning used by ViDS. This creates a real correctness risk: the benchmark may reward outputs that conform to a statistical face model rather than perceived motion fidelity. The reader's weakest assumption already identified this 3DMM-alignment concern, and the recommendation remains conditional acceptance pending verification. The concrete test with an independent motion metric and a different 3DMM estimator would settle whether the concern lands; if it does, the paper's claim should be more circumscribed.","tokens_in":16542,"tokens_out":7382,"duration_ms":65456,"concrete_test":"Re-run the VFHQ self- and cross-reenactment evaluations using an independent motion-fidelity measure not derived from 3DMM fitting—e.g., average facial keypoint (MediaPipe) distance or optical-flow endpoint error between generated and driving frames—and also recompute AED/APD with a different estimator (DECA or EMOCA). Then apply a paired bootstrap/permutation test to the metric differences among Ours, Wan-Animate, and HunyuanPortrait. If Ours does not remain ahead on the independent metrics and the AED/APD gaps shrink under EMOCA/DECA, the motion-control advantage is substantially an artifact of 3DMM-manifold overlap.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of more detailed expression/pose control rests on AED/APD/AKD plus a small user study. The load-bearing problem is that AED/APD are computed by Deep3DFaceRecon [18], a monocular 3DMM estimator, while ViDS's conditioning signal is itself a rendered 3DMM (Pixel3DMM/FLAME) normal map. Images generated from such geometry tend to stay near the 3DMM manifold, so a 3DMM-based evaluator will recover pose/expression parameters closer to the driver than a method that produces equally good but less parametric-conforming images. This makes the reported motion gaps vulnerable to estimator bias rather than purely perceptual fidelity. The paper's caveat that 'automated reconstruction can misalign with perceived quality' does not repair the numbers, because the substitute user study is 10 videos/40 participants with no significance testing. Independently of that bias, the tables do not consistently support the abstract: self-reenactment AED is 0.121 for Ours vs 0.118 for Wan-Animate (Table 1), and on Celeb-V-Text self AED is 0.189 vs HunyuanPortrait's 0.161 (Table 2); in cross-reenactment AED Ours is 0.298/0.306 while best baselines are 0.279/0.298. So the one metric specifically named 'expression distance' never favors ViDS, and self APD also favors Wan-Animate. A claim of 'more detailed and consistent expression and pose control' is therefore not anchored by the motion metrics shown.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes ViDS, a portrait animation method that uses Pixel3DMM/FLAME normal maps as dense geometric conditioning for a Wan-based video diffusion model. Identity shape is fixed from the reference image, driving pose/expression animates the mesh, and normal maps are concatenated with reference latents through a unified conditioning layer. Multi-CFG with separate identity/geometry/text guidance and autoregressive overlapping-window inference are used. Experiments on VFHQ and Celeb-V-Text compare to Follow-Your-Emoji, X-Portrait, HunyuanPortrait, and Wan-Animate, with ablations on conditioning signals, architecture, guidance scale, training size, and tracker. The paper reports best or second-best results on most reconstruction and identity metrics, and a user study favors ViDS on all dimensions.","tokens_in":16983,"tokens_out":6200,"duration_ms":50225,"significance":"If the evidence held, ViDS would be a useful and simple design: it injects dense 3DMM geometry into a pretrained video diffusion model with minimal architectural change, and it demonstrates the importance of tracking quality. The ablations are extensive and the training protocol is described in enough detail to reproduce. The main weakness is that the advertised advantage in expression/pose control is not consistently supported by the numerical motion metrics, and the motion metrics themselves share a 3DMM prior with the conditioning signal. The central claim is thus plausible but not conclusively demonstrated.","major_comments":[{"comment":"The abstract claims 'more detailed and consistent expression and pose control', but the only direct expression metric (AED) never favors ViDS. In Table 1 self-reenactment AED is 0.121 for Ours versus 0.113 (HunyuanPortrait), 0.118 (X-Portrait), and 0.118 (Wan-Animate); cross-reenactment AED is 0.298 versus 0.279 (HunyuanPortrait). In Table 2, Ours again has worse AED in both self (0.189 vs 0.161) and cross (0.306 vs 0.298) settings. This should be confronted directly: either temper the claim, show that AED is the wrong yardstick with a validated alternative, or provide a significance test showing the differences are not meaningful.","section":"Sec. 4.1, Tables 1 and 2"},{"comment":"AED and APD are computed with Deep3DFaceRecon [18], a monocular 3DMM estimator, while the conditioning signal is a rendered 3DMM (Pixel3DMM/FLAME) normal map. This creates a circularity concern: outputs that stay close to the 3DMM manifold may score better on these metrics even if they are not perceptually more faithful. The paper acknowledges the limitation and adds a user study, but the user study is too small (10 videos, 40 participants, no error bars or significance tests) to carry the load. I ask for (a) a non-3DMM motion metric (e.g., landmark-velocity or optical-flow-based expression/pose distance) or a second 3DMM estimator with a different topology, and (b) a per-item or paired analysis showing that automated motion metrics agree with the user-study 'expression/pose consistency' ratings.","section":"Sec. 4, Evaluation Protocol; Sec. 3.1"},{"comment":"No error bars, confidence intervals, or significance tests are reported anywhere. Several headline margins are very small: Table 1 self PSNR 21.19 vs 21.13, CSIM 0.879 vs 0.876, LPIPS 0.178 vs 0.187; Table 2 self PSNR 19.98 vs 19.01. With 50 test sequences, these differences may lie within noise. Please report per-metric standard deviations and paired significance tests (or bootstrap CIs) for at least the main comparison tables and for the key ablations in Table 4.","section":"Sec. 4, Tables 1-8"},{"comment":"The cross-reenactment protocol is under-specified: 'pairs the first frame of one test sequence as the reference identity with a different sequence as the driving video' does not state how many pairs are used, whether all ordered pairs are evaluated, or how the random pairing is fixed. Since cross-reenactment numbers depend strongly on the chosen pairs, please specify the exact evaluation set and report the number of generated videos.","section":"Sec. 4, Dataset and Evaluation Protocol"}],"minor_comments":[{"comment":"The column 'IQA' is never defined. Please state which image quality assessment is used and how it is computed.","section":"Tables 1, 2, 4, 6, 7"},{"comment":"The sentence 'Overall performance improves with scale despite metric-level fluctuations' is not supported by the table: several metrics degrade from 5K to 20K and improve only at 30K, while FID/FVD and IQA fluctuate non-monotonically.","section":"Table 7"},{"comment":"The user-study questionnaire uses a single 'Expression & Pose Consistency' rating. Because the paper's central claim concerns expression and pose separately, please either separate these two questions or justify merging them.","section":"Table 3"},{"comment":"The tracking-quality comparison reports only CSIM/AKD and not AED/APD; adding AED/APD would directly support the sentence 'Accurate tracking is therefore important for identity and geometry control.'","section":"Sec. 4.2, Table 8"},{"comment":"The autoregressive windowing description says 'nominal overlap of f frames' and 'stride h=s-f', but the pseudo-code/algorithm for tail alignment and pixel-space blending is not given. A short algorithm box or pseudo-code would help reproducibility.","section":"Sec. 3.3"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid technical core and extensive ablations, but the headline expression/pose claim is contradicted by the paper's own AED numbers. I think the authors can fix this within the manuscript's scope by re-analyzing the metrics, adding non-parametric motion metrics and uncertainty quantification, and softening the claim if needed. Note also that Pixel3DMM and SHeaP are from the same research group as several of the authors; the comparison is fair on its face, but the readers should be given independent motion metrics."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: ViDS is a genuinely well-engineered portrait-animation system, and the ablations are above average for the field. But the headline claim—more detailed and consistent expression control—is not supported by the paper's own motion metrics. The two things you should know: (1) the specific integration of Pixel3DMM normal maps as dense conditioning in a Wan video diffusion backbone, with unified channel concatenation and multi-CFG, is new and worth trying; (2) the AED/APD numbers, which are supposed to measure expression/pose accuracy, never favor ViDS in the self-reenactment tables, and on cross-reenactment AED ViDS is the worst of the five.\n\nWhat's good: The design is coherent. Keeping the reference identity mesh fixed while transferring driving pose/expression is a sensible disentanglement. The comparison of conditioning representations (normal maps vs UV map vs gray mesh vs landmarks) and the tracking-quality sensitivity table are useful contributions. The user study shows a clear perceptual preference (4.55 vs 3.84 next best), and the limitations section is honest about tracking sensitivity, slow inference, and lack of relighting.\n\nSoft spots: The evaluative circularity is real. AED/APD come from Deep3DFaceRecon, a monocular 3DMM estimator, and ViDS's control signal is itself a rendered 3DMM normal map. Any method that produces images staying close to the FLAME manifold will get artificially good AED/APD; baselines that produce equally good but less parametric-conforming images will look worse on those metrics. The paper acknowledges this and adds a user study, but the study is only 10 videos/40 participants with no significance testing. More importantly, even taking the metrics at face value, ViDS does not win on expression distance anywhere: VFHQ self AED 0.121 vs 0.118 (Wan-Animate) and 0.113 (HunyuanPortrait); cross AED 0.298 vs best 0.279; Celeb-V-Text self AED 0.189 vs 0.161. So the abstract overclaims. There are also no error bars or significance tests on any table, and the 15K dataset-size ablation shows an AED outlier (0.238 vs 0.121-0.129 for adjacent sizes), which suggests the numbers are noisy. No code or model release, so independent verification is limited.\n\nBottom line: The paper is a useful systems contribution for researchers working on portrait animation and video diffusion conditioning. It deserves a serious referee, but the review should ask for a cleaner motion-evaluation protocol (e.g., independent landmark-based or human-annotated expression metrics, a larger user study with statistical tests) and a reworded abstract that matches the evidence. I'd treat it as a conditional accept and would not desk-reject.","headline":"Solid, well-ablated systems paper; the dense 3DMM-normal-map conditioning is a real integration, but the motion metrics don't back the abstract's expression-control claim.","tokens_in":17445,"tokens_out":4504,"would_cite":true,"duration_ms":35866,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A video diffusion model can act as a neural shader for portrait animation when conditioned on dense, pixel-aligned 3D face normal maps rather than sparse landmarks or implicit motion latents.","keywords":["portrait animation","video diffusion","3D morphable model","normal map conditioning","identity preservation","face reenactment","autoregressive video generation","classifier-free guidance"],"falsifier":"Run the VFHQ self- and cross-reenactment evaluation while progressively corrupting or replacing the 3DMM tracking (e.g., jittering pose parameters or substituting SHeaP), and check whether identity and geometry metrics degrade as sharply as Table 8 suggests; if a corrupted tracking signal still yields strong CSIM and AKD, the conditioning is not doing the claimed work. A second decisive check is to re-evaluate all methods with a non-3DMM motion-fidelity metric or independent human ratings on expression transfer, since AED/APD share the paper's geometric prior.","tokens_in":16482,"feed_emoji":"🎭","tokens_out":7051,"duration_ms":55267,"temperature":0.7,"pith_summary":"ViDS sets out to prove that a video diffusion model can act as a neural shader for portrait animation when it is given dense geometric guidance from a 3D Morphable Model tracker. The core proposal is to render pixel-aligned normal maps from an identity-frozen 3D face mesh animated by a driving video, and concatenate them with the reference image and noisy latents into a pretrained video diffusion transformer. The authors argue this yields finer and more consistent expression and pose control than landmark-based or implicit-latent diffusion baselines, while preserving reference identity and appearance — including non-3DMM regions such as inner mouth and hair. The claim is supported by self- and cross-reenactment benchmarks on VFHQ and Celeb-V-Text, a user study, and ablations that isolate the conditioning signal and tracking quality.","feed_headline":"Dense 3D face tracking sharpens portrait animation","feed_subtitle":"Pixel-aligned 3D face normal maps give a video diffusion model sharper expression and pose transfer with identity kept.","key_machinery":"The central object is the rendered pixel-aligned normal-map sequence from an animated 3DMM mesh. Unlike low-dimensional 3DMM parameters or sparse landmarks, these maps carry dense local surface orientation that is lighting-agnostic and frame-aligned with the reference image; freezing the reference identity shape separates identity from pose and expression. This geometry channel is injected into a pretrained video diffusion transformer via a unified channel-concatenation layer (reference latents, normal-map latents, noisy video latents), and three separate classifier-free guidance branches (identity, geometry, text) are combined into one velocity field. Long sequences are produced by an autor","core_discovery":"ViDS claims that accurate monocular 3DMM tracking — specifically the Pixel3DMM tracker — converts expression and pose transfer into a shading problem. An identity-specific mesh is reconstructed from one reference image, then animated with a driving video's pose and expression parameters while shape parameters are frozen to prevent identity leakage; the animated mesh is rendered as normal maps and fed, together with the reference image and a text prompt, into a video diffusion transformer. The study reports that this dense geometric conditioning outperforms sparse-landmark and implicit-latent methods on most metrics in self- and cross-reenactment, with a particular advantage in identity prese","pith_inferences":["The paper's own limitation list (tracking sensitivity, slow autoregressive inference, no relighting, limited control beyond 3DMM regions) points to the clearest next steps: more robust tracking and faster sampling would expand the method's practical range.","Because the reported motion metrics (AED/APD) are computed with a 3DMM estimator, a re-ranking of methods using a different estimator or purely perceptual motion judgments would test whether the advantage is genuine or partly an artifact of shared geometric priors.","A testable extension: swapping the 3DMM tracker at inference for a stronger one should improve the same metrics on the same benchmarks, providing a direct way to measure how much headroom remains in tracking quality."],"forward_implications":["Dense geometric conditioning becomes a workable alternative to landmarks or implicit latents for one-shot portrait animation.","Improvements in monocular 3D face tracking quality should translate directly into finer expression and pose transfer, since the normal maps are the sole motion channel.","Identity leakage from the driving video is reduced by freezing the reference identity's 3DMM shape parameters, which should help cross-identity and in-the-wild reenactment.","The autoregressive overlapping-window scheme extends a pretrained video diffusion model beyond its native temporal window while reducing boundary discontinuities.","The method retains photorealistic synthesis of regions the 3DMM does not model, such as inner mouth and hair, because the diffusion prior still generates those details."],"fun_headline_variants":["3D face tracking turns portrait animation into a shading task","Expression transfer becomes shading with 3D face tracking","Video diffusion shader uses 3D face tracking for lifelike animation","ViDS: diffusion shading with 3D face tracking"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole pipeline leans on the assumption that Pixel3DMM tracking is accurate enough that its rendered normal maps faithfully capture the driving video's pose and expression — the paper's own conclusion lists tracking sensitivity as a limitation, and the ablation shows that swapping in SHeaP tracking sharply degrades identity and geometry metrics.","fun_headline_variants_meta":{"raw":{"variants":["3D face tracking turns portrait animation into a shading task","Expression transfer becomes shading with 3D face tracking","Video diffusion shader uses 3D face tracking for lifelike animation","ViDS: diffusion shading with 3D face tracking"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001022,"raw_usage":{"total_tokens":4120,"prompt_tokens":691,"completion_tokens":3429,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":435,"completion_tokens_details":{"reasoning_tokens":3372}},"tokens_in":435,"tokens_out":3429,"duration_ms":20928,"temperature":1.0,"reasoning_tokens":3372,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T22:57:01.017580+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the VFHQ self- and cross-reenactment evaluation while progressively corrupting or replacing the 3DMM tracking (e.g., jittering pose parameters or substituting SHeaP), and check whether identity and geometry metrics degrade as sharply as Table 8 suggests; if a corrupted tracking signal still yields strong CSIM and AKD, the conditioning is not doing the claimed work. A second decisive check is to re-evaluate all methods with a non-3DMM motion-fidelity metric or independent human ratings on expression transfer, since AED/APD share the paper's geometric prior.","supporting_citations":[],"review_version":1}