{"id":"78ad1c4e-44de-4271-9110-bbb8d115fcdc","arxiv_id":"2508.09973","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"PERSONA creates a personalized 3D avatar from one image by using diffusion-generated pose-rich videos to train a 3D Gaussian avatar with balanced sampling and geometry-weighted optimization.","lead":"PERSONA turns a single photo of a person into an animatable 3D avatar by first generating videos of that person in many poses with a diffusion model, then training a 3D Gaussian avatar on those videos. The result is an avatar that keeps the person's identity and also shows natural clothing motion when posed in new ways.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Need explicit disjointness of diffusion-training motion pool from NeuMan/X-Humans test poses; otherwise reported pose-deformation gains may reflect exposure to test poses.","rationale":"The paper's method is interesting and its qualitative results, ablations, and real-time rendering are credible. The reader's weakest assumption about diffusion-generated identity and geometry drift is real, and the paper's own limitations section acknowledges the lack of fine-grained cloth wrinkles and texture inconsistencies. However, the most load-bearing issue for the quantitative claim is the unspecified relationship between the target motion pool used for generating training videos and the test poses on NeuMan and X-Humans. If the same motion sources are used, the avatar is trained on the test pose distribution, which would inflate the pose-deformation gains relative to single-image baselines that do not see those poses. This is not an accusation of misconduct; it is a missing transparency statement about a critical evaluation detail. The paper explicitly says the qualitative videos are \"different from our training set\" but never makes the analogous statement for the quantitative benchmarks. The concrete test—inspecting the pose distributions or asking for the source list—would settle the concern quickly. If disjointness is confirmed, the central claim survives; if not, the numerical comparison is weakened. Since the reader already issued a conditional verdict, I recommend keeping that verdict unchanged while adding this specific condition to the list.","tokens_in":16509,"tokens_out":5963,"duration_ms":69011,"concrete_test":"Ask the authors for the list of source videos or SMPL-X pose sequences used to generate the ~1K training frames per test subject. Then, for each NeuMan and X-Humans test sequence, compute a pose-distance (e.g., per-joint angle L2 after Procrustes alignment, or nearest-neighbor search in pose space) between the training target poses and the test-frame poses. If any training pose lies within a small threshold (e.g., 5° mean joint-angle error) of any test pose—or if any source video comes from the same recording—the quantitative comparison is confounded. As a sensitivity check, retrain with a held-out motion set that deliberately excludes the test sequences and report the metric delta; a drop of more than ~1 dB PSNR would indicate that the reported advantage is partly test-pose memorization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—outperforming single-image methods on NeuMan/X-Humans—depends on evaluating generalization to novel poses. Section 4.2 builds the training set from target 3D poses \"extracted in advance from public videos using the ExAvatar fitting process,\" but the paper never states that these sources are disjoint from the benchmark test sequences. The only explicit disjointness statement is for the qualitative dance videos (\"different from our training set,\" Sec 7.1); no equivalent guarantee is given for Tables 2 and 3. If the target-motion pool includes the test motions (or near-identical poses), PERSONA is optimized on generated frames at the evaluation poses, while single-image baselines such as AniGS and LHM are not—so the measured gains in pose-driven deformation could reflect train/test pose overlap rather than a learned deformation model. This is a transparency gap in the evaluation protocol, not an accusation. The paper must either specify the motion source and its disjointness or make the source available to verify.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"PERSONA proposes a single-image whole-body avatar pipeline that combines SMPL-X/3DGS with diffusion-generated pose-rich training videos. The paper generates training videos with MimicMotion using target 3D poses extracted from public videos, then optimizes an ExAvatar-style Gaussian avatar using balanced sampling and geometry-weighted optimization. Experiments on NeuMan and X-Humans report state-of-the-art results among single-image methods, with ablations, a user study, runtime comparisons, and a limitations section. The central idea is to obtain pose-driven deformations without per-subject pose-rich video capture.","tokens_in":16767,"tokens_out":9789,"duration_ms":97927,"significance":"If the results hold, PERSONA is a practical and scalable contribution: it transfers pose-driven deformation knowledge from a diffusion-based animator into an explicit 3D avatar while preserving identity through sampling and geometry-weighted losses. The paper ships strong evidence in the form of quantitative comparisons on two benchmarks using official implementations, multi-component ablations, a 40-participant user study, and candid limitation statements. However, two evaluation-hygiene issues currently temper confidence in the main quantitative claim: the target motion pool is not shown to be disjoint from the benchmark test poses, and the loss weights are selected on the NeuMan test set. These are fixable protocol gaps rather than flaws in the core derivation.","major_comments":[{"comment":"The paper does not establish that the target motion pool used for diffusion-generated training videos is disjoint from the NeuMan and X-Humans test sequences used in Tables 2 and 3. Section 4.2 states that target 3D poses are \"extracted in advance from public videos using the ExAvatar fitting process,\" and Section 7.1 gives an explicit disjointness guarantee only for the qualitative in-the-wild dance videos (\"different from our training set\"). If the motion pool includes poses from the NeuMan/X-Humans test sets, PERSONA is optimized on generated frames at evaluation poses while single-image baselines are not, so the reported pose-deformation gains could partly reflect exposure rather than generalization. Please state the source of the motion pool, its overlap with the benchmark test sequences, and either release the pool or rerun with a provably disjoint pool.","section":"Sec. 4.2 / Sec. 7.1"},{"comment":"Table S2 reports loss-weight tuning directly on the NeuMan test set (\"Effect of loss weights ... on the NeuMan test set\"), with the chosen row marked ours. Table 3 then reports NeuMan results under this configuration. Thus the NeuMan comparison is not a fully held-out evaluation for the image-loss weight, and the reported 29.20 dB may be optimistically selected. Please select hyperparameters on a validation split or report sensitivity on both benchmarks with a fixed, pre-registered configuration. The balanced-sampling ratio deserves the same treatment: it is currently supported only by the qualitative Figure S8.","section":"Sec. 7.3 / Table S2"},{"comment":"The quantitative comparisons are reported without variance or the number of runs, although the pipeline involves stochastic diffusion-generated training videos (Sec. 4.1) and stochastic optimization. Some margins are modest, e.g., 0.80 dB over AniGS on X-Humans 00028 and less than 1 dB over the no-deformation ablation on NeuMan. It is therefore unclear whether the ranking is stable. Please report mean and standard deviation over at least three independent generation/training runs, or justify why the variance is negligible.","section":"Tables 2 and 3"}],"minor_comments":[{"comment":"Several figure labels contain untranslated Korean characters or garbled text (e.g., \"정국 127\" in Fig. 1 and similar artifacts in later figures). The camera-ready version should use English labels throughout.","section":"Fig. 1, Fig. 6, Fig. S2, Fig. S8"},{"comment":"The phrase \"we regularize these regions using separate RGBs\" is not defined in the architecture description. Please clarify whether these are additional optimizable color features, a separate rendered color branch, or simply a masking of the image loss.","section":"Sec. 5.1"},{"comment":"The baseline \"Ours wo. pose-driven deform.\" is not precisely specified. State exactly which modules are removed (mean-offset MLPs, triplane conditioning, geometry-weighted losses, or all of these) so the ablation is reproducible.","section":"Sec. 7.2 / Fig. 9"},{"comment":"The statement \"All comparisons exclude background pixels\" should specify how the foreground mask is obtained and whether the same mask is applied to every method. This is needed for a fair comparison of the reported PSNR/SSIM/LPIPS values.","section":"Sec. 7.1"},{"comment":"The generator comparison would be more informative if the paper stated whether the same target pose set and the same number of generated frames were used for all generators. Otherwise the small differences may reflect motion content rather than generator quality.","section":"Table S4"}],"recommendation":"major_revision","confidential_remarks":"The core idea is original and the experimental package is otherwise strong. The main risk is evaluation hygiene: the motion-pool/test overlap must be disclosed and the loss weights should not be selected on the test set. If the authors can provide a clear disjointness statement or release the motion pool and rerun with validation-based tuning, the paper would be suitable for acceptance. I do not see a circularity or novelty problem; ExAvatar is properly cited as prior work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"PERSONA deserves a serious referee. The core idea is genuinely new: instead of trying to perfect a single-image diffusion animator, it uses a diffusion model (MimicMotion) to generate pose-rich training videos of the input subject, then optimizes a 3D Gaussian avatar (built on ExAvatar) against those videos. The two introduced pieces—balanced sampling (oversampling the input image to fight identity drift) and geometry-weighted optimization (down-weighting image loss, up-weighting mask/depth/normal consistency when training the deformation MLPs)—both address real problems that would otherwise sink the pipeline. Ablations are informative, and comparisons are run with official implementations, which counts for a lot.\n\nThe paper is also honest about what it can't do: no dynamics, no fine wrinkles, no relighting, and an hour of preprocessing. That's the right tone.\n\nThe soft spots are evaluation-protocol issues rather than method flaws. The stress-test concern is legitimate: the training motions are 'extracted in advance from public videos' (Sec. 4.2), and unlike the qualitative dance videos, the paper never states that those public videos are disjoint from the NeuMan and X-Humans test sequences. If the motion pool includes test poses, PERSONA has effectively seen the evaluation poses during training (through generated frames), while the feed-forward baselines have not. That would inflate the pose-deformation gains in Tables 2 and 3. I don't see evidence of a deliberate leak—the motion sources are probably external—but the paper must state disjointness explicitly and ideally release the motion list. Second, the loss weights for geometry-weighted optimization are selected via an ablation on the NeuMan test set (Table S2). That is test-set tuning, which makes the headline numbers slightly optimistic. Third, there are no error bars or multi-run variance anywhere. Minor, but worth a sentence.\n\nThe representation is standard ExAvatar-style hybrid Gaussians with MLP mean offsets; nothing exotic, and nothing wrong. The dependence on diffusion-generated supervision is the real limitation, but it is the point of the paper, not a hidden flaw.\n\nBottom line: a genuine step for single-image avatar creation, with an evaluation far more thorough than most in this area. The disjointness statement and test-set tuning should be fixed in revision. I'd send it to peer review, and I'd cite it.","headline":"A genuinely new training-data pipeline for single-image avatars, with an honest evaluation that needs a clearer train/test pose-disjointness statement.","tokens_in":17186,"tokens_out":3129,"would_cite":true,"duration_ms":32058,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PERSONA builds a personalized whole-body 3D avatar with pose-driven clothing deformation from a single image by training on diffusion-generated videos.","keywords":["3D avatar","3D Gaussian splatting","SMPL-X","pose-driven deformation","diffusion video generation","identity preservation","single-image reconstruction","whole-body animation"],"falsifier":"Take one photo of a person in a loose skirt; generate MimicMotion videos; build PERSONA; then drive the avatar through a real motion-captured sequence where the skirt should flare. Compare the rendered skirt mask to a real video of the same person in the same motion: if the mask IoU does not improve over the diffusion generator's own frames, or if identity embedding distance to the input photo exceeds the generator's drift, the central claim of pose-driven deformation with preserved identity is falsified.","tokens_in":16442,"feed_emoji":"🧍","tokens_out":7517,"duration_ms":77158,"temperature":0.7,"pith_summary":"PERSONA combines the two dominant avatar-building strategies: 3D optimization, which preserves identity but needs pose-rich video, and diffusion-based animation, which learns deformations but loses identity. The paper's central proposal is to use a diffusion animator to synthesize pose-rich training videos from a single input image, then optimize a 3D Gaussian avatar on that footage. Two correctives make this transfer work: balanced sampling oversamples the input image to hold identity, and geometry-weighted optimization down-weights unreliable image textures in favor of stable geometry maps. If correct, this gives a scalable route to personalized, animatable whole-body avatars with natural cloth deformation from one casual photo, and the paper reports that it outperforms all existing single-image methods on NeuMan and X-Humans benchmarks.","feed_headline":"One photo becomes a 3D avatar whose clothes move with the body","feed_subtitle":"AI-generated pose-rich videos train a Gaussian avatar, preserving identity and sharp cloth deformation from a single image.","key_machinery":"The central object is a hybrid surface-mesh/3D-Gaussian body model anchored to SMPL-X: each template vertex carries an isotropic 3D Gaussian, and pose-driven deformation is produced by MLPs that take triplane canonical features plus the 3D poses of only 4-ring neighboring joints and output mean offsets (translations) to Gaussian positions before LBS animation and Mip-Splatting rendering. The mean-offset MLPs are the mechanism that generates non-rigid cloth movement; balanced sampling and geometry-weighted optimization are the two correctives that keep the MLPs learning from reliable identity and geometry signals rather than diffusion artifacts.","core_discovery":"The core claim is that a personalized whole-body 3D avatar with pose-driven non-rigid deformations can be obtained from a single image by optimizing a hybrid SMPL-X/3D-Gaussian avatar against diffusion-generated pose-rich videos. The paper argues that this requires two mechanisms: balanced sampling, which oversamples the input image and uses Sobel-detected seam boundaries plus albedo supervision to prevent identity drift and baked-in shadows; and geometry-weighted optimization, which sets low image-loss weights and high geometry-loss weights (masks, depth, normals, part segmentations) because geometry remains reliable where generated textures are inconsistent. It further claims that modeling","pith_inferences":["The framework's ceiling is set by the diffusion generator: a generator with stronger identity preservation would shrink the benefit of balanced sampling, while a generator producing 3D-consistent multi-view output could relax the geometry-weighted loss and preserve fine wrinkles.","Because only Gaussian means are shifted, the method trades away relighting, dynamic cloth/hair motion, and fine wrinkles; adding separate garment and hair layers that also update scales and colors is the natural next step.","The same 'synthesize pose-rich training data, then optimize a 3D representation' recipe should transfer to other single-image articulated 3D tasks, such as animals or deformable objects, where pose-varied footage is the bottleneck.","A practical test of the identity-preservation claim: run the pipeline with animators of different identity fidelity and measure how the avatar's identity distance to the input scales, quantifying how much of the gain is balanced sampling versus generator quality."],"forward_implications":["A single, casually captured photo becomes a fully animatable whole-body avatar, eliminating per-subject multi-view, 3D-scan, or pose-diverse video capture.","Pose-driven deformations like cloth lifting with raised arms are learned explicitly, avoiding the baked-in input deformations seen in prior single-image 3D methods.","Identity (face, clothing patterns) is preserved across novel poses better than the underlying diffusion animator, thanks to balanced sampling.","Rendering stays sharp in novel poses because deformation uses only mean offsets, and geometry supervision anchors optimization where textures are unreliable.","The pipeline renders in real time (about 25.6 fps on an A6000) after roughly one hour of video generation plus 30 minutes of avatar optimization."],"supporting_citations":[{"why":"Supplies the hybrid surface-mesh/3D-Gaussian representation, Laplacian regularization, and the SMPL-X fitting process used for target poses.","marker":"[37]"},{"why":"MimicMotion is the diffusion-based animator that generates the pose-rich training videos from the input image.","marker":"[66]"},{"why":"Provides the parametric whole-body model (identity shape, 3D pose, facial expression, and skinning) that the avatar is anchored to.","marker":"[40]"},{"why":"Mip-Splatting is the rendering technique used for alias-free 3D Gaussian splatting.","marker":"[64]"},{"why":"Sapiens supplies the geometry estimators (normal maps, part segmentations) used as stable supervision in geometry-weighted optimization.","marker":"[22]"},{"why":"SAM provides the binary masks used for geometry supervision.","marker":"[23]"},{"why":"Albedo images from ordinal-shading decomposition are used to prevent baked-in shadows during balanced-sampling supervision.","marker":"[5]"},{"why":"AniGS acts as the single-image 3D baseline that bakes input-specific deformations; used for quantitative and qualitative comparison.","marker":"[45]"},{"why":"NeuMan provides an in-the-wild dataset and evaluation protocol for rendering quality comparisons.","marker":"[19]"},{"why":"X-Humans provides a whole-body motion dataset used for evaluation and ablations.","marker":"[52]"}],"fun_headline_variants":["One photo, full-body 3D avatar with pose-driven cloth","Single-image avatar learns non-rigid cloth from generated videos","PERSONA: one image, personalized 3D avatar with realistic cloth","From one snapshot to animatable 3D human with pose deformations","Hybrid AI avatar: single image, no pose videos, cloth moves right"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The diffusion-generated videos must preserve both the identity and the pose-dependent appearance (cloth deformation, geometry) of the input subject closely enough that optimizing an avatar against them transfers real deformation behavior instead of generator artifacts.","fun_headline_variants_meta":{"raw":{"variants":["One photo, full-body 3D avatar with pose-driven cloth","Single-image avatar learns non-rigid cloth from generated videos","PERSONA: one image, personalized 3D avatar with realistic cloth","From one snapshot to animatable 3D human with pose deformations","Hybrid AI avatar: single image, no pose videos, cloth moves right"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000161,"raw_usage":{"total_tokens":1068,"prompt_tokens":735,"completion_tokens":333,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":479,"completion_tokens_details":{"reasoning_tokens":238}},"tokens_in":479,"tokens_out":333,"duration_ms":4104,"temperature":1.0,"reasoning_tokens":238,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:39:57.522672+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take one photo of a person in a loose skirt; generate MimicMotion videos; build PERSONA; then drive the avatar through a real motion-captured sequence where the skirt should flare. Compare the rendered skirt mask to a real video of the same person in the same motion: if the mask IoU does not improve over the diffusion generator's own frames, or if identity embedding distance to the input photo exceeds the generator's drift, the central claim of pose-driven deformation with preserved identity is falsified.","supporting_citations":[{"cited_title":"Segment any- thing","cited_arxiv_id":null,"evidence_quote":"SAM provides the binary masks used for geometry supervision."},{"cited_title":"Expressive whole-body 3D gaussian avatar","cited_arxiv_id":null,"evidence_quote":"Supplies the hybrid surface-mesh/3D-Gaussian representation, Laplacian regularization, and the SMPL-X fitting process used for target poses."},{"cited_title":"Mimicmo- tion: High-quality human motion video generation with confidence-aware pose guidance","cited_arxiv_id":null,"evidence_quote":"MimicMotion is the diffusion-based animator that generates the pose-rich training videos from the input image."},{"cited_title":"Expressive body capture: 3D hands, face, and body from a single image","cited_arxiv_id":null,"evidence_quote":"Provides the parametric whole-body model (identity shape, 3D pose, facial expression, and skinning) that the avatar is anchored to."},{"cited_title":"Mip-Splatting: Alias-free 3D gaussian splatting","cited_arxiv_id":null,"evidence_quote":"Mip-Splatting is the rendering technique used for alias-free 3D Gaussian splatting."},{"cited_title":"Sapiens: Foundation for human vision mod- els","cited_arxiv_id":null,"evidence_quote":"Sapiens supplies the geometry estimators (normal maps, part segmentations) used as stable supervision in geometry-weighted optimization."},{"cited_title":"Intrinsic image decomposi- tion via ordinal shading","cited_arxiv_id":null,"evidence_quote":"Albedo images from ordinal-shading decomposition are used to prevent baked-in shadows during balanced-sampling supervision."},{"cited_title":"AniGS: Animatable gaussian avatar from a single image with inconsistent gaussian reconstruction","cited_arxiv_id":null,"evidence_quote":"AniGS acts as the single-image 3D baseline that bakes input-specific deformations; used for quantitative and qualitative comparison."},{"cited_title":"NeuMan: Neural human radiance field from a single video","cited_arxiv_id":null,"evidence_quote":"NeuMan provides an in-the-wild dataset and evaluation protocol for rendering quality comparisons."},{"cited_title":"X- Avatar: Expressive human avatars","cited_arxiv_id":null,"evidence_quote":"X-Humans provides a whole-body motion dataset used for evaluation and ablations."}],"review_version":1}