{"id":"1ac1c421-b94f-4841-a419-1534d8f822e3","arxiv_id":"2509.06573","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A hybrid skeletal-diffusion pipeline animates hand-drawn characters from a single image with 3D guidance, inpainting coarse renders and injecting secondary dynamics via latent blending and hair-body separation.","lead":"This paper builds a hybrid animation system that renders a hand-drawn character with 3D skeletal motion, then uses a fine-tuned video diffusion model to redraw the frames with richer detail and secondary motion. It matters because it attacks the long-standing trade-off between geometric consistency and expressive dynamics when animating stylized drawings.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim of 'natural secondary dynamics' and quantitative SOTA superiority rests on static texture metrics (Table 1) with negligible differences and an unreported user study; no motion-aware evaluation supports the headline.","rationale":"The reader's verdict (CONDITIONAL) is reasonable: the proposed hybrid pipeline is plausible and the components are well described, but the evidence for the central claim is not yet sufficient. My stress-test identifies a load-bearing concern that is somewhat different from the reader's weakest_assumption: rather than the fidelity of the 3D proxy, the most consequential gap is the absence of any motion-sensitive evaluation of the claimed 'natural secondary dynamics' and quantitative SOTA superiority. The only quantitative table (Table 1) uses static texture metrics, reports negligible differences without error bars, and excludes the main diffusion baselines; the claimed user study is never described. The paper's own limitations (§6) acknowledge back-view long-hair failures, which supports the reader's proxy-fidelity concern but is a narrower issue. Because the central contribution (SDI) is about dynamics, the evaluation gap is load-bearing. However, the manuscript is not internally inconsistent, and the qualitative results and ablations suggest the system may well work; the appropriate remedy is to require the missing evaluation, not to reject the work outright. Hence UNCHANGED (remaining CONDITIONAL), with partial agreement with the reader because the same broad call is made but for a different primary reason.","tokens_in":15822,"tokens_out":6416,"duration_ms":62353,"concrete_test":"Run a pre-registered forced-choice perceptual user study on a held-out set of 20 characters and 5 motions, comparing Ours, DrawingSpinUp, UniAnimate*, UniAnimate, and MikuDance on two questions: 'which animation has more natural hair/clothing (secondary) motion?' and 'which better preserves the original drawing style?'. With 30 naive participants and paired bootstrap significance (e.g., 95% CI excluding 0.5), report the preference proportions and the full protocol. If Ours is not significantly preferred over DrawingSpinUp and UniAnimate* on secondary motion while maintaining style, the paper's headline claims should be weakened accordingly; additionally report the per-clip distribution of the Table 1 metrics to check whether the FID gap is within noise.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.4 (Table 1) is the only quantitative evidence for the abstract's claim that the method 'outperforms state-of-the-art methods in both quantitative and qualitative evaluations.' The three metrics (LPIPS, FID, CLIP) are all computed between generated frames and the reference image; they measure per-frame texture fidelity, not motion quality, temporal consistency, pose accuracy, or the secondary dynamics that the SDI module is introduced to provide. The reported differences are tiny (LPIPS 0.1733 vs 0.1734; FID 152.9 vs 157.7; CLIP 0.9030 vs 0.8964), no error bars or significance tests are given, and the main diffusion baselines (UniAnimate, AnimateAnyone, MikuDance) are excluded from the quantitative comparison. A 'perceptual user study' is claimed in the Introduction, but no protocol, sample size, stimuli, or statistics appear anywhere in the manuscript. The centrality of SDI (§3.5) and its ablation (§4.5.3) are argued qualitatively with selected frames; since no metric measures dynamics, the central claim that the hybrid pipeline conveys 'natural secondary dynamics' better than alternatives is not empirically established by the presented evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a hybrid animation system for hand-drawn characters that combines skeletal animation with video diffusion priors. A 3D proxy is reconstructed from a single drawing via Wonder3D and Mixamo, a Hair Layering Modeling (HLM) step separates hair from the body for long-hair characters, and the rigged model is retargeted to target motions to render pose, mask, and coarse color guidance sequences. A domain-adapted diffusion model, built on UniAnimate and fine-tuned on a synthesized drawing-animation dataset, refines these coarse frames. Secondary Dynamics Injection (SDI) blends latent estimates from the domain-adapted and pre-trained models during early denoising steps, followed by a re-denoising phase with Poisson-blended coarse guidance. The paper claims superior performance over both skeletal and diffusion-based state-of-the-art methods in quantitative and qualitative evaluations, and it reports ablations for HLM, the coarse prior encoder, spatial-layer tuning, and SDI components and hyperparameters.","tokens_in":16162,"tokens_out":3825,"duration_ms":38428,"significance":"If the claims are substantiated, the hybrid skeletal-plus-diffusion pipeline would be a useful contribution to stylized character animation: it directly addresses the geometric-consistency versus expressive-dynamics trade-off, and the HLM and SDI ideas are plausible and likely to interest the computer graphics community. The manuscript is also commendable for providing an anonymous code/data repository, detailed pipeline diagrams, and a broad set of ablations of the main components. However, the central quantitative claim of outperforming state-of-the-art methods is not currently supported by the evidence in the paper, and the key contribution (SDI) is evaluated only through selected qualitative frames. The paper is therefore best viewed as presenting a promising system whose headline claims require substantially stronger empirical support before publication.","major_comments":[{"comment":"The quantitative evidence for the abstract's claim that the method 'outperforms state-of-the-art methods in both quantitative and qualitative evaluations' is insufficient. The differences in Table 1 are extremely small (LPIPS 0.1733 vs. 0.1734, FID 152.90 vs. 157.75, CLIP 0.9030 vs. 0.8964), no error bars or significance tests are reported, and the metrics are computed per-frame against the reference image, measuring texture/semantic fidelity rather than motion quality, temporal consistency, pose accuracy, or secondary dynamics. The paper should add error bars across evaluation clips, significance tests, and motion-aware or temporal-consistency metrics (e.g., warping error, flow consistency, or a well-specified user study) before claiming quantitative superiority.","section":"Section 4.4, Table 1"},{"comment":"The main diffusion baselines are excluded from the quantitative comparison. The paper states that UniAnimate, AnimateAnyone, and MikuDance may keep reference texture and are therefore omitted, but these are the state-of-the-art methods the abstract claims to outperform. The comparison is restricted to DrawingSpinUp and UniAnimate*, the latter being fine-tuned on the same synthesized dataset used to train the proposed method. This setup biases the comparison in favor of the proposed system, and it does not establish superiority over the pose-controllable diffusion baselines. The authors should either include those baselines with a protocol that controls for the texture-preservation issue, or clearly delimit the quantitative claim to the compared methods.","section":"Section 4.4"},{"comment":"The central contribution, Secondary Dynamics Injection, is validated only through selected visual frames in Figures 13 and 14. Since the paper's headline claim is about 'natural secondary dynamics,' the lack of any quantitative or statistically tested evaluation of dynamics is load-bearing. The claim in Section 6 that the authors 'observed for the first time that different denoising steps are closely associated with distinct types of motion' is likewise supported only by a qualitative inspection of decoded latents in Figure 7. The paper needs either a quantitative dynamics evaluation, a complete user-study report with protocol and statistics, or a substantial softening of these claims.","section":"Sections 4.5.3 and 4.5.4"},{"comment":"There is a same-source dependency that is not fully acknowledged or tested. The training set is synthesized by DrawingSpinUp, which is also one of the two quantitative baselines and a prior method by the same group; the evaluation set contains characters from external drawing datasets but the ground-truth videos are again generated by DrawingSpinUp. This makes the comparison against DrawingSpinUp partly a test of how well the model reproduces the outputs on which it was trained, and does not establish generalization to real hand-drawn animations or to other animation pipelines. The authors should clarify the exact overlap between training and evaluation sources and provide at least one evaluation on data or annotations independent of the training pipeline.","section":"Section 4.1 and Section 4.4"},{"comment":"The system's robustness depends on the quality of the single-image 3D reconstruction and auto-rigging, but the paper provides no systematic analysis of failure rates or manual effort. The method requires manual hair-body segmentation maps for long-hair characters and relies on Mixamo auto-rigging, and Section 6 admits that back-view long-hair regions remain problematic (Fig. 18b). Given that HLM is claimed to address the main failure case of prior skeletal methods, the paper should quantify how often HLM succeeds, how much manual correction is needed, and how sensitive the final results are to reconstruction/rigging errors. Without this, the generality of the central claim is not established.","section":"Sections 3.2, 3.3, and 6"},{"comment":"The Introduction states that 'the results of comprehensive experiments and a perceptual user study' demonstrate the system's performance, but no user study appears in the manuscript. There is no protocol, number of participants, stimuli, task description, or statistical analysis. Since this sentence is part of the paper's evidence for superiority, the authors must either include the full user study or remove the claim. A qualitative figure comparison alone is not a substitute for a reported user study.","section":"Section 1 and Section 4.4"}],"minor_comments":[{"comment":"The notation in Eq. (5) mixes set operations on segmentation maps (∪, ∩) with element-wise multiplication of implicit fields (⊙); please clarify how binary masks are converted to the spatial domain of the implicit fields and how the 'back' mask is derived.","section":"Section 3.2, Eq. (5)"},{"comment":"The phrase 'manually-processed hair-body segmentation maps' should state explicitly that the user must provide these maps for long-hair characters and should indicate the typical effort involved, as this affects reproducibility.","section":"Section 3.2"},{"comment":"The statement that 'when τ1 < 0.7, the impact on motion distribution remains minimal' appears to contradict the earlier observation that β affects runtime; the relationship between β, motion content, and runtime should be stated more precisely.","section":"Section 3.5.3"},{"comment":"The metric descriptions should specify how per-frame LPIPS/FID/CLIP scores are aggregated across frames and clips, and why LPIPS is described as 'texture consistency' when it compares generated frames to the reference image.","section":"Section 4.4"},{"comment":"The axis labels in Figure 7 are unclear; please label the denoising timestep or percentage explicitly and state which model's latent estimates are being visualized.","section":"Figure 7 and Section 3.5.1"},{"comment":"The implementation details state that the model 'randomly decided whether to apply a mask' and also 'enhanced mask diversity'; please clarify the exact masking schedule and whether masks are applied to training videos consistently.","section":"Section 4.2"},{"comment":"The reference for Marching Cubes contains a typo ('Alghrithm'); please correct it.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is likely to be of interest to the SIGGRAPH/TOG audience, but the gap between the abstract's strong claims and the reported evidence is substantial. The same-source dependency between training data, baseline, and evaluation deserves particular scrutiny during revision. I would encourage the editor to require the authors to either provide a properly reported quantitative evaluation with error bars and full baselines, or to substantially rewrite the claims to match the qualitative evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely hybrid system with two new mechanisms that address real failure modes, and the qualitative story is mostly convincing. But the central claim — that it produces natural secondary dynamics better than existing methods — is not backed by the evaluation. Table 1 reports only static per-frame texture scores (LPIPS/FID/CLIP), with gaps like 0.1733 vs 0.1734, no error bars, no significance tests, and no motion-aware metric. The perceptual user study promised in the Introduction is absent from the manuscript.\n\nWhat's actually new: the SDI strategy — blending latent estimates from the adapted and pretrained diffusion models only inside user-specified masks during a middle denoising window, then re-denoising with Poisson-blended coarse frames — is a real mechanism, and the ablation shows it does something visible. The HLM hair/body separation via segmentation maps is a practical fix for a real artifact of single-mesh low-poly reconstruction. Fine-tuning only spatial layers while freezing temporal layers is a sensible choice given the small dataset, and the inpainting formulation of coarse-render enhancement is reasonable.\n\nThe soft spots are mostly in the evaluation. The same-source dependency is real: the training video is synthesized from DrawingSpinUp, and DrawingSpinUp is also a comparison baseline. This doesn't make the work circular, because the final output is not derived from that ground truth alone, but it makes the comparison against DrawingSpinUp more like an ablation of their own prior pipeline than a neutral SOTA comparison. Excluding UniAnimate, AnimateAnyone, and MikuDance from Table 1 is defensible if they fail on pose, but then the quantitative comparison is only against two skeletal baselines plus a fine-tuned UniAnimate. The \"observed for the first time\" claim about denoising steps is overstated; the phenomenon of early structure and late detail is known. The code and data are promised in an anonymous repository but not verifiable from the submission.\n\nTo the authors' credit, the limitations are stated honestly: inverted poses and back-view long hair remain problematic. That is the right kind of candor.\n\nWho this is for: graphics and animation researchers working on hand-drawn character animation, or on applying video diffusion to stylized content. It deserves a serious referee, but with a clear expectation of major revision: either add motion-aware evaluation (temporal consistency, pose accuracy, a real user study) or tone down the abstract. I would not desk-reject it.","headline":"Plausible hybrid pipeline with real new modules, but the headline claim of natural secondary dynamics is supported only by static texture metrics and an unreported user study.","tokens_in":16612,"tokens_out":1910,"would_cite":true,"duration_ms":19773,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A hybrid of skeletal animation and video diffusion animates a single hand-drawn picture in 3D while preserving style and adding natural hair and cloth motion.","keywords":["hand-drawn character animation","video diffusion models","skeletal animation","secondary motion","single-image 3D reconstruction","diffusion inpainting","character rigging","hair animation"],"falsifier":"Systematically degrade the 3D proxy quality, for example by feeding the pipeline drawings whose reconstruction has known limb-length or joint errors, and measure whether the final animation's pose accuracy and identity preservation degrade in step. If animation quality stays flat while proxy error grows, the 3D guidance is not actually carrying the result; likewise, a held-out set of long-haired back views should show the reported shoulder-texture errors if the hair-body separation is only partial.","tokens_in":15578,"feed_emoji":"🎬","tokens_out":8183,"duration_ms":62993,"temperature":0.7,"pith_summary":"The paper tries to settle a trade-off in hand-drawn character animation: skeletal rigging keeps a drawing's identity stable but renders hair and clothing stiffly, while video diffusion models produce lively motion but drift into real-human artifacts and broken contours. It claims that the two can be combined, in a split of labor, by first retargeting the drawing onto a rough 3D character, rendering coarse pose frames, and then treating the final output as an inpainting problem for a domain-adapted video diffusion model. The diffusion model is asked to refine only what skeletal animation does poorly, and a Secondary Dynamics Injection step blends in latent estimates from a pre-trained human-motion diffusion model so that masked regions such as hair ends and skirts gain realistic secondary dynamics. A Hair Layering Modeling step separates hair from body geometry before rigging so that long-haired characters do not deform unnaturally. If the claims hold, an animator needs only one drawing, a target 3D motion, and optionally a few masks to get stylized, editable animations with natural secondary motion.","feed_headline":"Hybrid rig-and-diffusion animates hand-drawn characters naturally","feed_subtitle":"Keeps the drawing's style while adding flowing hair and cloth motion, without simulation or manual rigging.","key_machinery":"The load-bearing object is the three-part guidance-and-refinement loop built around a rough 3D character: a single-image-to-3D module reconstructs a low-poly proxy with optional hair/body separation via segmentation maps; a domain-adapted video diffusion model conditioned on rendered coarse sequences refines appearance as an inpainting task; and Secondary Dynamics Injection blends v-prediction latent estimates $\\hat{z}^{v_\\theta}_{n,0}$ and $\\hat{z}^{u_\\theta}_{n,0}$ through masks $M^{\\text{SDI}}_{n,\\text{down}}$ during a middle denoising interval, followed by Poisson blending and a full re-denoise. The mechanism works because the denoising trajectory separates structure from detail and secondary motion, so the pre-trained model's real-human motion priors can be injected where secondary dynamics live without letting it redraw the whole character.","core_discovery":"The central claim is that geometric consistency and expressive dynamics are not competing requirements for hand-drawn characters; they can be assigned to different stages of one pipeline. Skeletal animation on a reconstructed 3D proxy supplies coarse guidance sequences that carry identity, viewpoint, and primary motion, and the diffusion stage is cast as inpainting so it redraws only user-marked regions rather than re-imagining the whole drawing. The paper's key observation is that in v-prediction denoising, early steps fix spatial structure and primary motion while later steps add secondary dynamics; the Secondary Dynamics Injection exploits this by blending the denoised latent estimates of the fine-tuned stylized model and the pre-trained real-human model over a middle denoising interval, then re-denoises from scratch with Poisson-blended inpainted coarse frames. Hair Layering Modeling separates hair from body in the implicit field using segmentation maps, so the reconstructed character has separate hair and body geometry instead of one fused low-poly mesh. The authors report that this outperforms both skeletal-only and diffusion-only baselines on their test set and that the components are each necessary in ablation.","pith_inferences":["Inference: the denoising-phase observation generalizes beyond hand-drawn characters, so any pose-conditioned video diffusion model could adopt the blend-and-redenoise trick to trade fidelity against injected dynamics in other stylized domains, such as furry creatures or loose clothing, that lack large training datasets.","Inference: the quality ceiling is set by the reconstructed 3D proxy, so as single-image-to-3D reconstruction improves, the same pipeline should improve without retraining; a testable prediction is that animation fidelity correlates with proxy reconstruction accuracy.","Inference: the SDI masks could be turned into an interactive 'dynamics brush' where the user paints where motion should be amplified and the system re-runs only the affected denoising phases, making animation editing a local, near-real-time operation.","Inference: Hair Layering Modeling could be extended to layered clothing by segmenting multiple garment layers in the implicit field rather than just hair versus body."],"forward_implications":["Given a single frontal drawing and a target 3D motion, the pipeline outputs a stylized 3D animation that preserves the drawing's identity, including long-hair characters that skeletal methods deform.","User-supplied masks can steer which regions receive secondary dynamics, and edits made to the reference frame propagate across the whole animation without re-rigging.","Because refinement is framed as inpainting, the system needs only a small stylized training set; fine-tuning spatial layers while freezing temporal layers preserves the motion prior and avoids catastrophic forgetting.","The method does not require physics simulation or manual multi-layered rigging, so it lowers the barrier for novices to animate single drawings."],"supporting_citations":[{"why":"Supplies the image-to-3D reconstruction and skeletal animation baseline; its outputs are also used as ground truth to fine-tune the stylized diffusion model.","marker":"[Zhou et al. 2024]"},{"why":"The base pose-controlled video diffusion architecture and pre-trained weights; it becomes both the domain-adapted model and the pre-trained model used for secondary-motion injection.","marker":"[Wang et al. 2024c]"},{"why":"The cross-domain single-image-to-multi-view module that produces the neural implicit field from which the 3D proxy is reconstructed.","marker":"[Long et al. 2024]"},{"why":"The auto-rigging service used to rig the reconstructed character and retarget the target 3D motion.","marker":"[Inc. 2023]"},{"why":"2D skeletal animation baseline and source of the Amateur Drawings Dataset used to build the training set.","marker":"[Smith et al. 2023]"},{"why":"Source of the SketchAnim Dataset used, together with the Amateur Drawings, to construct the small-scale stylized animation training set.","marker":"[Rai et al. 2024]"},{"why":"Provides the v-prediction parameterization used to estimate denoised latents that the SDI blending operates on.","marker":"[Salimans and Ho 2022]"},{"why":"Poisson blending is used to stitch the estimated video into the masked coarse frames before the final re-denoise.","marker":"[Pérez et al. 2003]"},{"why":"Supplies the lightweight STC-encoder architecture adapted as the coarse prior encoder that embeds masked coarse color and SDI masks.","marker":"[Wang et al. 2023]"}],"fun_headline_variants":["Diffusion adds natural motion to hand-drawn character rigs","Rigged skeletons guide diffusion for lively 2D animation","Hybrid rig-diffusion brings hair and cloth to cartoon characters","Inpainting-based diffusion animates hand-drawn characters naturally","Skeleton plus diffusion for expressive secondary dynamics in cartoons"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole pipeline leans on the reconstructed and auto-rigged low-poly 3D character being accurate enough that its rendered coarse frames are trustworthy guides for identity and pose; if single-image reconstruction or auto-rigging goes wrong, the diffusion refinement inherits the error, and the paper itself notes back-view long-hair regions still misbehave.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion adds natural motion to hand-drawn character rigs","Rigged skeletons guide diffusion for lively 2D animation","Hybrid rig-diffusion brings hair and cloth to cartoon characters","Inpainting-based diffusion animates hand-drawn characters naturally","Skeleton plus diffusion for expressive secondary dynamics in cartoons"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000545,"raw_usage":{"total_tokens":2643,"prompt_tokens":1014,"completion_tokens":1629,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":630,"completion_tokens_details":{"reasoning_tokens":1544}},"tokens_in":630,"tokens_out":1629,"duration_ms":11075,"temperature":1.0,"reasoning_tokens":1544,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:14:56.188771+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Systematically degrade the 3D proxy quality, for example by feeding the pipeline drawings whose reconstruction has known limb-length or joint errors, and measure whether the final animation's pose accuracy and identity preservation degrade in step. If animation quality stays flat while proxy error grows, the 3D guidance is not actually carrying the result; likewise, a held-out set of long-haired back views should show the reported shoulder-texture errors if the hair-body separation is only partial.","supporting_citations":[],"review_version":2}