{"id":"0e87edf2-3444-41df-8cc2-6758fc89ea33","arxiv_id":"2505.02192","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"DualReal jointly trains identity and motion adapters for video customization, reporting a 21.7% CLIP-I and 31.8% DINO-I improvement over existing methods.","lead":"A team from USTC proposes DualReal, a training strategy that alternates identity and motion adapters in text-to-video personalization while using a controller to balance them across denoising steps, and reports large gains in identity fidelity metrics. The paper's headline numbers are potentially inflated by comparing against baselines running on different base video models.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Comparison does not isolate DualReal's contribution: published baselines use different backbones and fewer training steps, so the claimed 21.7%/31.8% gains and motion-metric wins may reflect base-model strength or compute rather than the method.","rationale":"The reader's weakest_assumption was the cross-backbone comparison; I agree that is a major threat and extend it: even the same-backbone baselines train for far fewer steps than DualReal, so the within-backbone gaps may be a training-budget effect rather than a method effect. The method itself is plausible, clearly described, and its ablations show consistent internal trends, so this is not a rejection. The appropriate remedy is to control for backbone and compute budget. I also note the headline percentage improvements do not match any obvious calculation from Table 1; subject-averaged percentages can legitimately differ from aggregate ratios, but the paper provides no per-subject data or formula, so that number needs a supplementary breakdown. The verdict therefore remains CONDITIONAL: accept if a matched-budget, matched-backbone comparison confirms the advantage; otherwise the central quantitative claim is unsupported.","tokens_in":14310,"tokens_out":13528,"duration_ms":156901,"concrete_test":"Re-run the CogVideoX-5B LoRA baseline and the CogVideoX-5B sequential full-parameter baseline with exactly 1,000 training steps, learning rate 1e-3, AdamW, and the same prompts and evaluation protocol as DualReal (per Supplementary 7.1); if the CLIP-I and DINO-I gaps over LoRA (0.204 and 0.265) shrink by more than half, the claimed gains are not attributable to the joint-training design.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is causal: adaptive joint training with gradient masking and stage/depth weighting produces the reported identity-fidelity and motion-quality gains. Table 1 cannot support that attribution. DualReal is evaluated on CogVideoX-5B, while MotionBooth uses LaVie-base and DreamVideo uses ModelScopeT2V (per Supplementary 7.1), so the two published baselines confound method with backbone. The same-backbone baselines are not matched in training budget: LoRA runs 600 total steps, full-parameter fine-tuning runs 330 steps, and DualReal runs 1,000 steps with a different learning rate (1e-3 vs 1e-4). The same-backbone gaps (CLIP-I 0.629 vs 0.425/0.521; DINO-I 0.551 vs 0.286/0.424) could therefore reflect additional optimization compute rather than the proposed architecture. In addition, the abstract's 'improves by 21.7% and 31.8%' is not reproducible from Table 1 under any obvious aggregation, so the magnitude of the claimed advantage is also unverified from the paper alone.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes DualReal, a framework for customized text-to-video generation that jointly trains identity and motion adapters on a Diffusion Transformer backbone (CogVideoX-5B). The method alternates identity- and motion-focused training steps with gradient masking to prevent cross-dimension leakage, and uses a StageBlender Controller to modulate the contribution of the two adapters as a function of denoising timestep and DiT block depth. The authors claim average improvements of 21.7% in CLIP-I and 31.8% in DINO-I over existing methods, and state that the approach achieves top performance on nearly all motion quality metrics. The paper includes quantitative comparisons in Table 1, component ablations in Tables 2 and 3, qualitative results, and supplementary implementation details.","tokens_in":14536,"tokens_out":3476,"duration_ms":49937,"significance":"If the central claim were established, the paper would make a useful contribution to video customization: the idea of alternating identity and motion training with gradient masking, combined with stage- and depth-dependent weighting, is a plausible mechanism for reducing identity-motion conflicts, and the paper provides a clearly described framework with detailed hyperparameters in the supplementary material. The work also assembles a benchmark with 50 identities and 21 motion sequences, which is a useful evaluation resource. However, the quantitative evidence as presented does not support the causal attribution of the reported gains to the proposed method, because the headline comparisons are confounded by base model and training budget, and the claimed average improvement percentages are not reproducible from the reported table. The central technical idea remains defensible, but the evaluation needs substantial strengthening before the paper's main claims can be accepted.","major_comments":[{"comment":"The main quantitative comparison does not isolate the proposed method. DualReal is evaluated on CogVideoX-5B, while MotionBooth uses LaVie-base and DreamVideo uses ModelScopeT2V, so the two published baselines confound method with backbone. The same-backbone baselines are also not matched in training budget: LoRA runs 600 total steps, full-parameter fine-tuning runs 330 steps, and DualReal runs 1,000 steps, with different learning rates (1e-3 vs 1e-4). The large same-backbone gaps in CLIP-I (0.629 vs 0.425/0.521) and DINO-I (0.551 vs 0.286/0.424) could therefore reflect additional optimization compute rather than the proposed architecture. The authors should either run baselines on the same backbone with matched training steps and learning-rate schedules, or clearly report the backbone and budget of every row in Table 1 and restrict causal claims to controlled comparisons.","section":"Sec. 4.1, Table 1, Supp. 7.1"},{"comment":"The claimed average improvements of 21.7% on CLIP-I and 31.8% on DINO-I are not reproducible from Table 1 under any standard aggregation. Averaging the four baselines gives (0.566+0.425+0.521+0.458)/4 = 0.4925 for CLIP-I, for which 0.629 is a 27.7% relative improvement, not 21.7%; for DINO-I the baseline average is 0.3758, and 0.551 is a 46.6% improvement, not 31.8%. The paper should state the exact formula used for the average improvement and report per-case or per-baseline numbers with error bars.","section":"Abstract and Table 1"},{"comment":"The 'lossless fusion' claim and the statement of 'top performance on nearly all motion metrics' are contradicted by the Dynamic Degree results. Relative to the benchmark average of 12.02, DualReal deviates by +2.94, whereas MotionBooth deviates by only -1.07, and DreamVideo deviates by -3.18, so DualReal is not the closest to the reference motion intensity. The authors acknowledge that their DD is 'not high,' but this undercuts the lossless-fusion language in the title and abstract. The paper should either soften the claim or provide a principled tolerance band within which a deviation is considered acceptable, with statistical support.","section":"Table 1, Dynamic Degree row"},{"comment":"The ablation studies report single runs with no variance or significance testing, and they are performed on a smaller subset whose size is not specified. Because identity and motion metrics are noisy for generative model evaluation, the observed differences (e.g., DINO-I 0.771 vs 0.766 for removing weight groups in Table 2) may not be meaningful. The authors should report standard deviations over multiple seeds or, at minimum, the number of cases in the ablation subset and the per-case metric distribution.","section":"Tables 2 and 3"}],"minor_comments":[{"comment":"Temporal Flickering is described as a mean absolute difference between adjacent frames, which normally increases with flicker, yet the table labels it with an up-arrow as a positive metric. The paper should clarify the sign convention and whether higher or lower is better.","section":"Sec. 4.1, Evaluation metrics"},{"comment":"The dimensions of W_down, W_up, and W_cond are not specified; providing these would make the adapter architecture and parameter counts fully reproducible.","section":"Sec. 3.2, Eq. (1)-(2)"},{"comment":"The mask notation is confusing: M is defined as Z*Mm + (1-Z)*Mi, but the conditions in Eq. (6) use l for the motion mask and k for the identity mask without defining the indexing over layers or parameter blocks. Please clarify that the masks are applied elementwise to the adapter parameters and specify how the per-block grouping from the StageBlender Controller interacts with this masking.","section":"Sec. 3.2, Eq. (4)-(6)"},{"comment":"The text 'Deprecated prompts' in the figure appears to be a typo and should be replaced with a meaningful label such as 'diverse prompts' or 'depicted prompts'.","section":"Figure 3"},{"comment":"The row labeled 'w/o Weight Groups' in Table 2 is numerically identical to the n=1 row in Table 3; the paper should state explicitly that these settings are the same, or explain why the same configuration appears in both tables.","section":"Tables 2 and 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is on a timely topic and the method is plausible, but the headline quantitative claims are not currently supported by the experimental design. The confound between method and backbone/budget in Table 1 is the main issue; it is fixable within the scope of a revision by adding controlled same-backbone, matched-budget baselines and by recomputing or rephrasing the average improvement claims. I would also encourage the editor to ask for explicit clarification of the Dynamic Degree deviation because the phrase 'lossless fusion' is contradicted by the paper's own numbers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: DualReal is a sensible, clearly-described method for training identity and motion adapters jointly in a DiT video model. It alternates adapter updates with gradient masking and uses a timestep/depth-dependent controller to weight adapter contributions. That combination is not in DreamVideo or MotionBooth; each piece is borrowed, but the joint-training framing is a reasonable step forward and the ablations suggest each component matters.\n\nThe paper's real contribution is the recipe, not a new theory. The writing is clear, the framework is easy to follow, and the ablations (Tables 2-3, Fig. 6-7) give a plausible internal story: removing the alternating scheme or the controller degrades identity and motion metrics. The visual analysis of controller weights showing depth-dependent identity/motion specialization is a nice touch.\n\nThe soft spots are real but not fatal. The headline numbers are not trustworthy as attribution. Table 1 compares DualReal on CogVideoX-5B against MotionBooth on LaVie and DreamVideo on ModelScopeT2V. That confounds method with backbone. The same-backbone baselines (LoRA, full FT) use fewer training steps (600/330 vs 1000) and, for full FT, a different learning rate. So the reported CLIP-I 0.629 vs 0.425/0.521 and DINO-I 0.551 vs 0.286/0.424 could be partly extra compute, not the architecture. The abstract's \"21.7% and 31.8% average improvement\" cannot be reproduced from Table 1 under any obvious averaging, which is sloppy. Also \"lossless fusion\" is overclaiming when the Dynamic Degree deviates by +2.94 from the 12.02 reference; that's not a collapse, but it isn't lossless either. No error bars or significance tests anywhere, so we have no sense of run-to-run variance for a method that depends visibly on hyperparameters (gamma, learning rate, steps).\n\nNone of this kills the idea. The joint-training mechanism is plausible and the internal ablations support it. But the paper needs same-backbone, matched-budget comparisons and a defensible aggregation for the headline gains before the empirical claim is established.\n\nWho is this for: people working on personalized video generation and adapter-based customization. It deserves a serious peer review, though I'd send it back for major revisions, not desk reject.","headline":"A plausible joint-training recipe for identity-motion video customization, but the reported gains are not cleanly attributable to the method because the head-to-head comparisons confound backbone, compute, and training budget.","tokens_in":15042,"tokens_out":2553,"would_cite":false,"duration_ms":28957,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Jointly training identity and motion adapters in one alternating loop, with gradient masking and stage/depth-dependent weighting, fuses both dimensions without loss in customized text-to-video generation, lifting CLIP-I by 21.7% and…","keywords":["video customization","text-to-video generation","identity consistency","motion consistency","joint training","diffusion transformer","gradient masking","adaptive weighting"],"falsifier":"Re-run all baselines (DreamVideo, MotionBooth, LoRA, full fine-tuning) on the same CogVideoX-5B backbone with matched training budgets; if DualReal's identity-similarity margins shrink to near zero or its motion metrics no longer lead, the joint-training claim would be falsified.","tokens_in":14116,"feed_emoji":"🎬","tokens_out":6831,"duration_ms":69654,"temperature":0.7,"pith_summary":"DualReal addresses a specific failure mode in customized text-to-video generation: when identity and motion are tuned separately, improving one degrades the other. The paper argues that the two dimensions are intrinsically coupled—stable identity restricts possible motions, and motion trajectories force identity changes—so a single coordinated training procedure should outperform isolated adapter tuning. Its central claim is that adaptive joint training, alternating identity and motion steps while masking gradients to prevent knowledge leakage and weighting contributions by denoising stage and network depth, fuses both patterns without loss. The reported result is a 21.7% average gain in CLIP-I and 31.8% in DINO-I over existing methods, with top or near-top scores on motion consistency metrics. A sympathetic reader would care because it suggests a general recipe for training multi-dimension customization in diffusion models.","feed_headline":"One shared training loop fuses identity and motion in custom videos","feed_subtitle":"DualReal alternates identity and motion steps with gradient masking, lifting CLIP-I by 21.7% and DINO-I by 31.8%.","key_machinery":"The central mechanism is a pair of complementary units. Dual-aware Adaptation alternates the training stage with a binary selector variable $Z$, applies gradient masks $M_m$ and $M_i$ that activate only the motion or identity adapter parameters, and keeps the frozen adapter as a latent regularizer during the other's forward pass. StageBlender Controller is a gated MLP conditioned on timestep embeddings and pooled text-visual features; it outputs softmax-weighted groups $\\omega^{(1)}\\dots\\omega^{(n)}$ that scale each DiT block's motion contribution, with identity scaled by $1-\\omega^i$, through the residual update $\\hat{f}^i_{\\text{out}} = \\omega^i f^i_{\\text{mo}} + (1-\\omega^i) f^i_{\\text{id}} + f^i_{\\text{dit}}$. The two together let the model vary the identity/motion balance across denoising steps and network depths.","core_discovery":"The paper's central claim is that the 'isolated customized paradigm'—training identity and motion adapters separately, then blending at inference—systematically degrades both dimensions because it ignores their interdependence and applies uniform optimization across all denoising steps. DualReal instead trains the two adapters in one shared loop: at each step a binary switch chooses identity or motion data, the active adapter is guided by the frozen prior of the other dimension, and a gradient mask blocks updates to the non-training adapter, preventing cross-dimension knowledge leakage. A StageBlender Controller then generates per-block scaling weights conditioned on the denoising timestep and fused text-visual features, so identity receives fine-grained attention in shallow, early blocks and motion receives growing weight in the deepest block as denoising proceeds. The paper reports that this arrangement improves identity similarity scores by 21.7% on CLIP-I and 31.8% on DINO-I on average across a constructed benchmark of 50 identities and 21 motion sequences, while matching or exceeding baselines on temporal consistency, motion smoothness, and flickering.","pith_inferences":["A direct test of the method's contribution would be to re-run DreamVideo and MotionBooth on the exact CogVideoX-5B backbone used for DualReal; the paper's quantitative comparison keeps different base models for the baselines, so a same-backbone comparison would isolate the joint-training effect from the base model.","The identity-weight curve in Figure 7—identity weight rising with denoising step except in the deepest block, which goes the opposite way—suggests a learned curriculum that maps naturally onto other multi-concept generation tasks, where attribute-specific schedules could be read off from similar controller analyses.","The framework implicitly predicts that any two customizable dimensions with asymmetric spatial or temporal structure will benefit from stage-dependent weighting; a cheap test would be applying the same controller to style-and-layout or subject-and-environment customization."],"forward_implications":["If DualReal's joint training is right, then the standard practice of training adapters for different attributes separately and merging at inference should be revisited; joint alternating training with gradient masking offers an alternative that avoids mutual degradation.","The denoising-stage and depth-dependent weighting suggests that identity and motion occupy different timescales in the diffusion process; one could schedule other attribute pairs, such as style and viewpoint, in the same way.","The gradient-masking regularization gives a concrete low-cost mechanism to prevent catastrophic interference in multi-task adapter training.","The constructed benchmark of 50 identities, 21 motion sequences, and 50 prompts per case could become a common testbed for identity-motion customization if other groups adopt it."],"supporting_citations":[{"why":"Supplies the isolated-paradigm baseline (separate identity and motion adapters blended at inference) that DualReal is designed to outperform, and the source of the bottleneck-adapter and conditional linear map design used in equations (1) and (2).","marker":"[52]"},{"why":"Baseline that preserves motion capability by injecting random videos during identity training; represents the prior approach to mitigating cross-dimension interference that gradient masking replaces.","marker":"[53]"},{"why":"The DiT backbone on which DualReal is built and the reference for the expert-transformer features and Adaptive LayerNorm used in the StageBlender Controller.","marker":"[56]"},{"why":"Baseline method that trains separate low-rank identity and motion adapters and fuses their parameters at inference.","marker":"[18]"},{"why":"The fine-tuning paradigm used to build the sequential full-parameter CogVideoX-5B baseline.","marker":"[38]"},{"why":"Source of the Motion Smoothness, Temporal Flickering, and Dynamic Degree metrics used to report motion quality.","marker":"[22]"},{"why":"Defines the DINO-I identity-similarity metric.","marker":"[4]"},{"why":"Defines the CLIP-T and CLIP-I similarity metrics.","marker":"[35]"}],"fun_headline_variants":["Joint training fuses identity and motion for better custom videos","One loop trains both identity and motion, boosting scores","DualReal unifies identity and motion in a single training step","Adaptive joint training ends identity-motion conflicts in videos","Fused identity and motion via adaptive joint training"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported improvements assume that comparing DualReal (trained on CogVideoX-5B) against baselines that run on different base models (MotionBooth on LaVie-base, DreamVideo on ModelScopeT2V) isolates the effect of the proposed training method; if the base model drives most of the metric gap, the gains would not be attributable to joint training.","fun_headline_variants_meta":{"raw":{"variants":["Joint training fuses identity and motion for better custom videos","One loop trains both identity and motion, boosting scores","DualReal unifies identity and motion in a single training step","Adaptive joint training ends identity-motion conflicts in videos","Fused identity and motion via adaptive joint training"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000164,"raw_usage":{"total_tokens":1282,"prompt_tokens":1017,"completion_tokens":265,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":633,"completion_tokens_details":{"reasoning_tokens":199}},"tokens_in":633,"tokens_out":265,"duration_ms":3399,"temperature":1.0,"reasoning_tokens":199,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:57:56.936799+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run all baselines (DreamVideo, MotionBooth, LoRA, full fine-tuning) on the same CogVideoX-5B backbone with matched training budgets; if DualReal's identity-similarity margins shrink to near zero or its motion metrics no longer lead, the joint-training claim would be falsified.","supporting_citations":[{"cited_title":"Dreamvideo: Composing your dream videos with customized subject and motion","cited_arxiv_id":null,"evidence_quote":"Supplies the isolated-paradigm baseline (separate identity and motion adapters blended at inference) that DualReal is designed to outperform, and the source of the bottleneck-adapter and conditional linear map design used in equations (1) and (2)."},{"cited_title":"Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation","cited_arxiv_id":null,"evidence_quote":"The fine-tuning paradigm used to build the sequential full-parameter CogVideoX-5B baseline."}],"review_version":1}