{"id":"7a6767e9-806c-4b57-8570-dc7d05e75388","arxiv_id":"2411.13536","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Swapping SDS for likelihood distillation and adding SVD rank weighting plus mirror and grid score gradients improves identity preservation in text-guided 3D head stylization with PanoHead.","lead":"This paper fine-tunes the PanoHead 3D face generator using likelihood-based diffusion distillation, SVD rank weighting, and mirror and grid score gradients to stylize heads in 360-degree views while preserving identity. If it works, game and VR developers can turn a single photo into a consistent stylized 3D head that still looks like the same person.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"FID/CLIP stylization metrics are compared against a reference distribution generated by each method's own diffusion teacher, so the quantitative stylization-quality claims in Table 1 are not cross-method comparable.","rationale":"The central claim is that the proposed distillation recipe (LD + rank weighting + mirror/grid gradients) improves both identity preservation and stylization quality relative to prior 3D domain adaptation. The identity half is supported by ArcFace ID, ΔD, and a user study, and the qualitative figures are consistent with that. The stylization-quality half, however, rests almost entirely on FID and CLIP numbers that are computed against a per-method reference distribution generated by each method's own diffusion teacher. Because the reference distribution changes from method to method, a lower FID can simply reflect better imitation of the teacher rather than better stylization. This is a textbook case of a self-referential benchmark. The ablation study has an additional, related confound: the LD variants use RV v5.1 while the initial SDS baseline may use a different checkpoint, so the improvement cannot be cleanly attributed to LD. These issues do not invalidate the identity-preservation evidence, but they do invalidate the quantitative stylization-quality comparison that the abstract highlights. Since the paper is otherwise detailed, the method is plausible, and the identity claim has independent support, a conditional acceptance with a re-run of the evaluation against a common reference is the appropriate outcome. This matches the reader's verdict, so no change is needed.","tokens_in":17897,"tokens_out":8204,"duration_ms":77641,"concrete_test":"Recompute Table 1 FID and CLIP for all methods against a single fixed reference distribution, e.g., generate the edited ground-truth images with Stable Diffusion v1.5 (or another held-out text-to-image model) for every method instead of each method's own checkpoint. If the ranking of methods changes, the reported stylization-quality improvements are an artifact of the self-referential benchmark. In addition, in the Table 2 ablation, run the DiffusionGAN3D/SDS baseline with RV v5.1 as the denoiser (same as the LD runs) to isolate the effect of the LD objective from the checkpoint change.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing concern is the validity of the stylization-quality numbers. In the Supplementary ('Quantitative scores'), the ground-truth edited distribution is constructed for each method by adding noise at t=25 to the same ~100 images and denoising for 50 steps with 'each baseline's diffusion checkpoints.' The paper's own pipeline uses RV v5.1 as the denoiser, while the baselines use their original checkpoints (e.g., SD1.5, SD2). Consequently, FID and CLIP in Table 1 measure how close each method's outputs are to a reference generated by that method's own teacher. A method that merely imitates its teacher will trivially score well; the numbers are not comparable across methods because the reference changes. This directly undermines the abstract's claim of 'substantial quantitative improvements' in stylization quality. The identity metrics (ArcFace ID, ΔD) and the user study are not affected by this circularity, so the identity-preservation half of the claim retains support. A second confound appears in the ablation (Table 2): the LD variants employ RV v5.1 while the [29] SDS baseline may use its original checkpoint, so the large ID jump from [29] to LD could be partly due to the teacher model, not the distillation objective.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a 3D head stylization method built on PanoHead that fine-tunes the generator using a negative log-likelihood distillation (LD) objective instead of SDS, together with rank-weighted SVD of score tensors, mirror-pose gradients, and multi-view grid denoising with a depth-conditioned ControlNet. The authors claim improved identity preservation (ArcFace ID, depth difference ΔD, and a user study) and improved stylization quality (FID, CLIP, KID) relative to several 3D domain-adaptation and 2D-3D editing baselines, on five style prompts. The central technical claims are that LD avoids the identity-collapse and over-smoothing observed with SDS, and that the mirror and grid extensions improve multi-view consistency without sacrificing identity.","tokens_in":18104,"tokens_out":7036,"duration_ms":74048,"significance":"If the identity-preservation claims hold, the paper makes a useful contribution: it demonstrates that an LD-style objective can be used for GAN domain adaptation with better identity retention than SDS, and the rank-weighting and multi-view distillation extensions are practical and well-ablated. The qualitative results, 360-degree visualizations, and user study are encouraging, and the method is described in enough detail to be reproducible. The main weakness is that the quantitative stylization-quality metrics are self-referential and are therefore not comparable across methods as presented; this issue must be corrected before the comparative conclusions can be fully accepted. The identity-related metrics (ID, ΔD) and the user study are not affected by this circularity.","major_comments":[{"comment":"The ground-truth edited distribution used for FID and CLIP is generated separately for each method by denoising with that method's own diffusion checkpoint, as stated in the Supplementary ('we take images, add noise with t = 25, and denoise with the style prompt using each baseline’s diffusion checkpoints for 50 steps'). Since the proposed method uses Realistic Vision v5.1 while StyleGANFusion, DiffusionGAN3D, and other baselines use their original SD1.5/SD2 checkpoints, Table 1's FID/CLIP values compare each method's outputs against a reference generated by that method's own teacher. A method that simply imitates its teacher would score well, and the numbers are not comparable across rows. This directly undermines the abstract's claim of 'substantial quantitative improvements' in stylization quality. The KID table in the Supplementary inherits the same problem. Please recompute the stylization metrics against a single fixed reference distribution (e.g., using one SD checkpoint for all methods) or, at minimum, restrict the cross-method claim to the identity metrics, which are not affected by this issue.","section":"Supplementary, 'Quantitative scores'; Table 1"},{"comment":"The ablation in Table 2 does not hold the diffusion teacher fixed. The row '[29]' uses SDS with DiffusionGAN3D's original checkpoint, while the rows 'LD + [63]' and onward use Realistic Vision v5.1, as stated in the Supplementary ('As the conditional denoiser for our method and the ablation study for showing the improvements upon [29], we employ RV v5.1'). The large ID improvement from 0.35 to 0.46 and the FID decrease from 147.19 to 116.88 are therefore not attributable solely to the change from SDS to LD; they may reflect the stronger teacher. Please re-run the [29] baseline with the same teacher (RV v5.1) used for the LD ablation rows, and re-report the deltas, so that the effect of the distillation objective is isolated.","section":"Table 2; Supplementary B (Implementation details)"}],"minor_comments":[{"comment":"The change-of-variables expression p(x_0) = p(x_t)|∂x_t/∂x_0|^{-1} = p(x_t)√ᾱ_t is dimensionally incorrect; the determinant should be ᾱ_t^{d/2} for a d-dimensional latent. The subsequent gradient formula is unaffected because the constant drops out under differentiation, but the equality as written is misleading.","section":"Supplementary Eq. (5)"},{"comment":"In the mirror-gradient derivation, the score ∇_{x_t} log p(x_t^π|y) appears in both terms of Eq. (13) even though the second term involves x_t^{π′}. Please state explicitly that, by the assumed symmetry, the score for the mirrored pose is evaluated on the mirrored tensor, or rewrite the indices so that the two terms are unambiguous.","section":"§3.2, Eqs. (13)-(14)"},{"comment":"The rank-weighting matrix W = diag(1, 0.75, 0.5, 0.25) is said to be set 'based on our empirical analysis' and used for all prompts. A short sensitivity study (e.g., different decay schedules) or a clearer justification for why these exact coefficients are robust would strengthen the generality claim.","section":"§3.2, 'Rank weighted score tensors'"},{"comment":"The statement that SDS is 'inherently mode-seeking' because of the subtraction of the ground-truth noise ϵ is stated without formal support. If retained, it should be presented as an empirical observation or supported by a reference, rather than as a derived property.","section":"§3.2, 'SDS vs LD on GANs'"}],"recommendation":"major_revision","confidential_remarks":"The core technical idea (using LD for GAN domain adaptation, with rank-weighted scores and mirror/grid extensions) is plausible and the identity-preservation evidence is reasonably strong. The main obstacle is the stylization-quality evaluation: both the cross-method comparison in Table 1 and the ablation in Table 2 conflate the method with the diffusion teacher. These are fixable within the manuscript's scope by using a common reference distribution and a common teacher in the ablation, so I recommend major revision rather than rejection. The paper is otherwise clearly written and the supplementary material is detailed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the identity-preservation half of this paper is credible; the stylization-quality half is not substantiated by the numbers. The FID/CLIP reference distribution is generated by each baseline's own diffusion checkpoint, so Table 1 measures how close each method is to its own teacher, not how well methods compare. That is a real flaw, not a nit.\n\nWhat's actually new: the rank-weighted SVD on VAE score channels is a genuine, simple trick that seems to control color artifacts. Combining LD (borrowed from PlacidDreamer) with mirror gradients and grid distillation for a 3D GAN is a sensible adaptation, and the paper is honest that LD is not theirs. The ablations are well structured and the qualitative figures look convincing, especially the accessory preservation (glasses, earrings). The ArcFace ID scores and the user study support the claim that identity is preserved better than the StyleGAN-based baselines.\n\nSoft spots, in order of severity. First, the stylization metrics. The supplementary describes building a per-method 'ground-truth' edited distribution by adding noise at t=25 and denoising with each baseline's diffusion checkpoint, then computing FID/CLIP against that. Since the method uses RV v5.1 and baselines use SD1.5/SD2, the numbers in Table 1 are not cross-method comparable. This doesn't sink the identity claims, but it undercuts the 'substantial quantitative improvements' in stylization quality. Second, the ablation may have a confounder: if the [29] baseline uses a different teacher than the LD rows, part of the ID jump could be teacher choice, not the distillation objective. The text is ambiguous; worth clarifying. Third, the diversity-preserving 3D baselines DATID-3D and PODIA-3D are cited but never compared, which is a noticeable omission for a paper about identity/diversity. Fourth, no error bars or significance tests anywhere, and no code. None of these make the central idea wrong; they make the evidence weaker than the abstract claims.\n\nWho this is for: anyone working on 3D-aware GAN domain adaptation or distillation from diffusion to one-step generators. The rank-weighting idea alone is worth a look.\n\nRecommendation: send it to review, but with a request for a fixed, non-circular evaluation—either a single reference model for all methods or a perceptual study—plus the missing baselines and variance reporting. The method deserves referee time; the current evaluation doesn't support the strong quantitative claims as written.","headline":"Plausible identity-preserving 3D head stylization with a genuinely neat rank-weighting trick, but the stylization-quality numbers are not cross-method comparable because each method is measured against its own diffusion teacher.","tokens_in":18703,"tokens_out":2844,"would_cite":true,"duration_ms":28593,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that PanoHead can be fine-tuned, using likelihood distillation and three gradient modifications, to stylize heads across 360 degrees while preserving each person's identity.","keywords":["3D head stylization","identity preservation","likelihood distillation","score distillation sampling","PanoHead","GAN domain adaptation","multi-view consistency","SVD rank weighting"],"falsifier":"Recompute FID/CLIP with a reference distribution generated independently of the distillation teacher, such as a second diffusion model or a fixed human-rated set of stylized portraits; if the margin over DiffusionGAN3D and StyleGANFusion shrinks or reverses, the quality gains rest on the benchmark's circularity. Separately, replace ArcFace with a different face-recognition embedding and re-measure identity similarity; if the gap disappears, the identity claim is ArcFace-specific.","tokens_in":17661,"feed_emoji":"🎨","tokens_out":5653,"duration_ms":59808,"temperature":0.7,"pith_summary":"This paper is trying to establish that the choice of distillation objective—not just extra regularization—determines whether text-guided 3D head stylization destroys identity. It shows that fine-tuning PanoHead with negative log-likelihood distillation (LD) rather than SDS, together with rank-weighted score tensors, mirror gradients, and grid-based multi-view distillation, yields stylized heads that remain recognizable as the same person from every viewing angle. If true, identity preservation becomes a built-in property of the distillation process rather than a separate loss that must be bolted on, and the approach should transfer to other 3D GANs and diffusion teachers.","feed_headline":"Identity survives 3D stylization via likelihood distillation","feed_subtitle":"Fine-tuning PanoHead with likelihood distillation instead of SDS keeps the original identity across every viewing angle.","key_machinery":"The central object is the distillation gradient used to fine-tune the PanoHead generator. LD replaces the SDS gradient, which subtracts the sampled noise epsilon, with the pure score estimate, so the update direction seeks higher likelihood rather than mode collapse. The SVD rank weighting decomposes the 4-channel score tensor and reweights its singular values with W = diag(1, 0.75, 0.5, 0.25) to keep the dominant stylization rank while suppressing the lower-rank color tint. Mirror gradients reuse the same score for yaw-symmetric poses and flip the backpropagated gradient, exploiting the head symmetry prior. Grid distillation forms a 2×2 grid of four render poses, denoises the grid jointly with a depth-conditioned ControlNet, and backpropagates the grid gradient before the SR network.","core_discovery":"The central claim is that a 3D-aware GAN like PanoHead can be adapted to a text-prompted style while preserving the input subject's identity by swapping SDS for LD, which is diversity-seeking rather than mode-seeking, and by adding three gradient-level modifications: re-weighting the SVD of the score tensor along the VAE channel dimension to suppress color artifacts, forcing cross-view consistency through mirrored gradients for symmetric poses and a 2×2 grid of renders passed to a depth-conditioned ControlNet, and routing grid gradients before the super-resolution network to avoid a resolution mismatch. The paper asserts that this recipe lets different input identities remain distinguishable after stylization, whereas prior methods such as DiffusionGAN3D, StyleGANFusion, StyleCLIP, and StyleGAN-NADA tend to produce similar outputs for different people.","pith_inferences":["The FID/CLIP gains should be re-checked against a reference set not produced by the same diffusion teacher; if the margin over baselines disappears, the stylization-quality advantage may be an artifact of a benchmark that favors the teacher's own outputs.","The mirror-gradient trick suggests a general recipe for any pose pair with a known equivariance map, so a rotation operator could extend the method to non-symmetric views and to full 3D objects beyond heads.","Because the diffusion model is frozen and only the GAN is adapted, the approach should port directly to newer, stronger 3D generators and diffusion teachers, provided the score tensor's channel dimension remains accessible for SVD weighting."],"forward_implications":["The same distillation recipe can be transferred to other 3D-aware GANs with a symmetric pose pair and a super-resolution stage, without per-prompt identity losses.","Identity preservation becomes a property of the objective, so less per-prompt tuning and fewer hand-crafted regularizers are needed compared with SDS-based baselines.","The rank-weighted score tensors provide a user-controllable dial for how much of the style's low-rank structure is applied, trading stylization strength against identity fidelity.","Distillation from diffusion models to one-step generators is improved by modeling cross-view dependencies instead of treating each render as independent, which the paper demonstrates through mirror and grid gradients."],"supporting_citations":[{"why":"Supplies the negative log-likelihood distillation objective and its derivation that the paper partially adapts for GAN fine-tuning.","marker":"[22]"},{"why":"PanoHead is the 3D-aware generator backbone that provides 360-degree renders and a two-stage super-resolution pipeline.","marker":"[4]"},{"why":"DiffusionGAN3D is the SDS-based baseline with relative-distance regularization that the paper builds on and compares against.","marker":"[29]"},{"why":"Defines Score Distillation Sampling, the objective that LD replaces and that the paper argues causes identity collapse.","marker":"[44]"},{"why":"ControlNet provides the depth-conditioned guidance used in grid and mirror distillation to keep geometry consistent across views.","marker":"[63]"},{"why":"The latent diffusion model and VAE embedding define the score estimation space that all distillation and rank weighting operate on.","marker":"[46]"}],"fun_headline_variants":["LD beats SDS for identity-preserving 3D head stylization","Multiview score distillation keeps 3D head identity intact","Swapping SDS for LD preserves face in 3D stylization","Identity-preserving 3D head stylization via likelihood distillation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that the method's stylization quality beats the baselines assumes that the FID and CLIP reference distribution produced by adding noise at timestep 25 to real images and denoising with each baseline's own diffusion checkpoint is a fair, method-neutral benchmark.","fun_headline_variants_meta":{"raw":{"variants":["LD beats SDS for identity-preserving 3D head stylization","Multiview score distillation keeps 3D head identity intact","Swapping SDS for LD preserves face in 3D stylization","Identity-preserving 3D head stylization via likelihood distillation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000275,"raw_usage":{"total_tokens":1616,"prompt_tokens":892,"completion_tokens":724,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":651}},"tokens_in":508,"tokens_out":724,"duration_ms":8263,"temperature":1.0,"reasoning_tokens":651,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:18:11.117461+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute FID/CLIP with a reference distribution generated independently of the distillation teacher, such as a second diffusion model or a fixed human-rated set of stylized portraits; if the margin over DiffusionGAN3D and StyleGANFusion shrinks or reverses, the quality gains rest on the benchmark's circularity. Separately, replace ArcFace with a different face-recognition embedding and re-measure identity similarity; if the gap disappears, the identity claim is ArcFace-specific.","supporting_citations":[{"cited_title":"PlacidDreamer: Advancing harmony in text-to-3D gen- eration, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the negative log-likelihood distillation objective and its derivation that the paper partially adapts for GAN fine-tuning."},{"cited_title":"Ogras, and Linjie Luo","cited_arxiv_id":null,"evidence_quote":"PanoHead is the 3D-aware generator backbone that provides 360-degree renders and a two-stage super-resolution pipeline."},{"cited_title":"DiffusionGAN3D: Boosting text-guided 3D generation and domain adaptation by combining 3D GANs and diffusion priors","cited_arxiv_id":null,"evidence_quote":"DiffusionGAN3D is the SDS-based baseline with relative-distance regularization that the paper builds on and compares against."},{"cited_title":"Barron, and Ben Milden- hall","cited_arxiv_id":null,"evidence_quote":"Defines Score Distillation Sampling, the objective that LD replaces and that the paper argues causes identity collapse."},{"cited_title":"Adding conditional control to text-to-image diffusion models","cited_arxiv_id":null,"evidence_quote":"ControlNet provides the depth-conditioned guidance used in grid and mirror distillation to keep geometry consistent across views."},{"cited_title":"High-resolution image synthesis with latent diffusion models","cited_arxiv_id":null,"evidence_quote":"The latent diffusion model and VAE embedding define the score estimation space that all distillation and rank weighting operate on."}],"review_version":1}