{"id":"ce56a9f6-93d5-4064-99ed-1ab47aa3d83a","arxiv_id":"2501.05369","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A single-network virtual try-on model with modality-specific normalization and shared attention matches or beats dual-network reference-based models on image and video try-on benchmarks.","lead":"MN-VTON replaces the extra reference network used in modern virtual try-on systems with a single diffusion model that handles text, garment image, and person video inputs through separate normalization steps, then shared attention. The paper reports higher image and video try-on quality, especially for details like logos and text, at lower or comparable parameter counts.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 1's LPIPS values are in inconsistent units (5.412 vs 0.0632), so the controlled evidence for choosing DiT and for lowering overhead is invalid; the central single-network claim lacks its key head-to-head support.","rationale":"I read the paper's strongest claim as a causal one: modality-specific normalization lets a single network replace a ReferenceNet, achieving comparable or better quality with lower overhead. The reader's weakest assumption about Eqs. (6)-(7) is reasonable, but the grouping choice is at least plausibly supported by the PCA visualization and ablation and does not, by itself, invalidate the single-network paradigm. A more immediately checkable and more decisive weak point is the controlled comparison in Table 1. That table is the only head-to-head, same-base-model evidence for the headline claim, and its metric column is internally inconsistent: values of 5.412 and 5.357 are orders of magnitude above every LPIPS value elsewhere in the paper, while the DiT row reports 0.0632. The table therefore cannot support the statement that Ours(DiT) significantly outperforms the dual-network baseline, nor does the parameter count support 'significantly reducing computational overhead' for the chosen DiT model. The verdict should remain conditional rather than move to rejection, because the issue is fixable by reporting correct units and compute measurements; if the clarified numbers reverse the Table 1 comparison, the central claim would lose its main controlled support. I partially agree with the reader because the modality-grouping assumption is a genuine risk, but I see the Table 1 inconsistency as the most load-bearing concern in the argument as written.","tokens_in":11342,"tokens_out":7020,"duration_ms":64585,"concrete_test":"Recompute Table 1 with consistent units and report both FID and LPIPS for all three rows. If the first two rows were LPIPS×100, normalized values make Ours(DiT) worse on LPIPS than Ours(UNet), so the architecture-selection claim fails. Also report inference FLOPs, peak memory, and latency for Reference-UNet, Ours(UNet), and Ours(DiT) at 512×384; if per-step compute of Ours(DiT) is not substantially below the dual-network baseline, the 'significantly reducing computational overhead' component of the central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on Table 1, the only controlled comparison between a dual-network ReferenceNet baseline and the proposed single-network variants. In that table, LPIPS is listed as 5.412 for Reference-UNet and 5.357 for Ours(UNet), but 0.0632 for Ours(DiT). Every other LPIPS value reported in the paper is between roughly 0.019 and 0.283 (Tables 2 and 3), so values of 5.4 cannot be standard LPIPS. If the first two rows are LPIPS multiplied by 100, the normalized values become 0.05412 and 0.05357 versus 0.0632 for Ours(DiT), meaning the DiT variant is actually worse on LPIPS than the UNet variants, contradicting the claim that DiT 'significantly outperforms' ReferenceNet. If the 5.4 values are actually FID, then the column is mislabeled and the DiT row lacks the required LPIPS entry. Either way, Table 1 cannot be used to justify the architecture selection. Additionally, the chosen Ours(DiT) has 1694.28M trainable parameters versus 1700.25M for Reference-UNet, so 'significantly reducing computational overhead' is not supported by the reported parameter counts; only the UNet variant halves parameters, and its advantage is marginal even under a favorable reading. Since the paper's strongest claim is that a single network matches or surpasses dual networks, and Table 1 is the only same-backbone, controlled evidence for that claim, this inconsistency is load-bearing.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MN-VTON, a single-network virtual try-on method for both images and videos. The core idea is a Modality-specific Normalization strategy: text tokens and visual tokens (garment image plus target image/video) are normalized and modulated separately with AdaLN-zero, while all tokens share the same attention layers in a diffusion Transformer built on CogVideoX. The authors argue that this addresses the three limitations they attribute to earlier single-network methods and removes the need for a dual-network ReferenceNet. They report quantitative comparisons on VITONHD and DressCode for image VTON and on VVT and VIVID for video VTON, together with an architecture-generalization experiment, ablations, and a user study, concluding that a single network can match or surpass dual-network approaches with reduced computational overhead.","tokens_in":11635,"tokens_out":6137,"duration_ms":58505,"significance":"If the claims are correct, the result is practically significant: it would challenge the current ReferenceNet-based dual-network consensus in VTON and offer a simpler, more scalable architecture for high-resolution image and video try-on. The paper is clearly written, the proposed normalization split is simple and easy to understand, and the experimental campaign covers four benchmarks plus a user study. The core derivation in Eqs. (4)-(7) is not circular and the approach is validated on held-out test sets. However, the central architecture-selection evidence and the efficiency claim currently rest on inconsistent metric reporting in Table 1, and the video comparison in Table 3 contains duplicated rows with conflicting values. These issues must be corrected before the significance of the single-network claim can be properly assessed. No code or weights are provided, which also limits reproducibility.","major_comments":[{"comment":"The LPIPS column mixes incompatible units: the entries 5.412 and 5.357 for Reference-UNet and Ours(UNet) are two orders of magnitude larger than every other LPIPS value in the paper, which ranges roughly from 0.019 to 0.283 in Tables 2 and 3. If the 5.4 entries are LPIPS multiplied by 100, then Ours(DiT) at 0.0632 is actually worse on LPIPS than both UNet variants, contradicting the statement that Ours(DiT) significantly outperforms ReferenceNet; if the 5.4 entries are FID, the column is mislabeled and the DiT row lacks an LPIPS value. Since this table is the only controlled ReferenceNet-versus-single-network comparison in the paper, the architecture selection and the central single-network claim depend on correcting these numbers, stating the metric definitions, and describing which split (validation vs. test) was used for model selection.","section":"§4.2, Table 1"},{"comment":"The claim that the method 'significantly reduces computational overhead' is not supported by the reported parameter counts. Ours(DiT) has 1694.28M trainable parameters versus 1700.25M for Reference-UNet, which is essentially the same; only the Ours(UNet) variant halves the parameter count. The paper should either restrict the efficiency claim to the UNet variant or report inference FLOPs, latency, or peak memory for the selected DiT configuration.","section":"§4.2 and §5"},{"comment":"The rows for VIVID and Ours are each duplicated with substantially different metric values (VIVID SSIM 0.949 vs. 0.8747; Ours SSIM 0.971 vs. 0.8879), and the text does not explain which numbers come from the original publications and which are the authors' re-evaluations under a common protocol. Without a clear statement distinguishing cited baselines from recomputed numbers, the video VTON comparison and the claimed improvements are ambiguous and should be clarified.","section":"Table 3"},{"comment":"The central design choice in Eqs. (6)-(7) is ablated only qualitatively; no quantitative metrics (SSIM, LPIPS, FVD) are reported for configurations (a)-(f). Given that the modality-specific grouping is the key novelty and the paper explicitly argues that grouping garment and target tokens while separating text is the most effective approach, a quantitative ablation table on VIVID or VITONHD is needed to support this claim.","section":"§4.4 and Fig. 7"}],"minor_comments":[{"comment":"The header 'VTIONHD' is a typo and should read 'VITONHD'.","section":"Table 2 caption"},{"comment":"References [36] and [37] are duplicate entries for the same VITON-DiT paper, and the in-text citation switches between [36] in Section 2 and [37] in Section 4.3; these should be merged.","section":"References [36] and [37]"},{"comment":"The text refers to supplementary material for the experiment on splitting each of the three modalities independently, but the arXiv v1 does not include supplementary text; please include the supplement or summarize the result in the main paper.","section":"§3.3"},{"comment":"The user study section reports preference percentages but does not specify the number of participants per study, the exact question format, or whether the preference differences are statistically significant; these details should be added.","section":"§4.5"},{"comment":"No code or pretrained weights are provided. For a method whose main contribution is an architectural design, releasing code and weights would substantially improve reproducibility and would also allow independent verification of the Table 1 and Table 3 metric issues.","section":"General"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Modality-specific normalization is a clean idea and the SOTA numbers are consistently strong, but the paper's only controlled head-to-head against a ReferenceNet is undermined by an LPIPS unit mess in Table 1, and the final DiT model uses essentially the same parameter count as the dual-network baseline, so the central 'single-network with lower cost' story is not currently supported.\n\nWhat's new: a diffusion transformer for VTON that splits tokens into text and visual (garment + target) groups, applies separate AdaLN-zero to each, then lets them attend together. That grouping is well motivated — garment and target are both images, text is different — and the ablations in Sec 4.4 and Fig 3 support it. Joint image-video training with position embedding interpolation is also a practical recipe. The reported gains over CatVTON, IDM-VTON, and others on four benchmarks are consistent and substantial. Credit where due: the component-level evidence is much stronger than in the average VTON paper.\n\nNow the soft spots. Table 1 is the only controlled comparison where the same backbone is used with and without a ReferenceNet. It lists LPIPS as 5.412 for Reference-UNet and 5.357 for Ours(UNet) — values about 100x larger than every other LPIPS in the paper, including 0.0632 for Ours(DiT). If the first two rows are LPIPS x 100, then Ours(UNet) (0.05357) is basically tied with Reference-UNet (0.05412), and Ours(DiT) (0.0632) is actually worse on that metric. That does not support the sentence in Sec 4.2 saying DiT 'significantly outperforms' ReferenceNet. Second, the trainable parameters for Ours(DiT) are 1694.28M vs 1700.25M for Reference-UNet — a 0.35% reduction, not a 'significant reduction in computational overhead.' That claim only holds for the UNet variant, which is not the one they use for the rest of the paper. Third, the final architecture is selected from Table 1 without a described held-out validation split, and there are no error bars or significance tests. Fourth, CatVTON is a single-network method and is cited, but the paper does not discuss it as such; the 'renaissance' framing therefore overstates the novelty of the single-network idea, even if the specific normalization approach is new.\n\nNet: the modality normalization and the joint training recipe are worth knowing; the SOTA table is likely broadly right; but the load-bearing controlled evidence is defective in its current form. A serious referee should ask for a corrected Table 1, per-metric confidence intervals, code, and a more honest efficiency comparison. If that comes back, this could be a solid CVPR-level contribution.","headline":"A clean normalization idea with consistently strong benchmark numbers, but the only controlled comparison against ReferenceNet is corrupted by an LPIPS unit error and the efficiency claim does not hold for the final DiT model.","tokens_in":12201,"tokens_out":2973,"would_cite":false,"duration_ms":25666,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single diffusion network with modality-specific normalization matches or beats dual-network virtual try-on on image and video benchmarks while cutting trainable parameters.","keywords":["virtual try-on","single-network paradigm","modality-specific normalization","diffusion transformer","video virtual try-on","ReferenceNet","garment fidelity","AdaLN-zero"],"falsifier":"Run the same MN-VTON pipeline with the alternative three-way split (text, garment, target each normalized independently) on VITONHD and VIVID test sets. If that configuration matches or beats the proposed two-way grouping on SSIM/LPIPS/VFID, the central claim that garment and target must share normalization collapses.","tokens_in":11079,"feed_emoji":"👗","tokens_out":4184,"duration_ms":35401,"temperature":0.7,"pith_summary":"The paper challenges the prevailing dual-network paradigm in virtual try-on, where a separate ReferenceNet extracts garment features to feed the main generation network. It proposes MN-VTON, a single diffusion transformer that handles text, garment image, and target person image or video by normalizing each modality separately and then letting all tokens share the same attention layers. The authors claim this single network matches or surpasses dual-network state-of-the-art methods on VITONHD, DressCode, VVT, and VIVID, while cutting trainable parameters roughly in half in the UNet configuration. If true, high-fidelity image and video try-on no longer requires the extra network, making higher resolution and longer video practical.","feed_headline":"Modality-specific normalization revives single-network try-on","feed_subtitle":"Text and visual tokens get separate normalization, letting one network top image and video try-on benchmarks.","key_machinery":"The central object is Modality-Specific Normalization: the network's token features are partitioned by input modality, with text tokens normalized by one AdaLN-zero layer and garment and target image/video tokens concatenated and normalized by another, before all tokens enter shared self-attention. This implements the split-and-fusion idea of Eq. (5) with a modality-based splitting rule (Eqs. (6)–(7)). The design preserves the relative relationships within a multimodal distribution and keeps garment and target information throughout the layers, so a single network can play the role that ReferenceNet played in prior work.","core_discovery":"MN-VTON shows that the supposed need for a dedicated ReferenceNet is an artifact of how features are normalized and fused, not a fundamental requirement. By splitting the shared token sequence into a text stream and a visual stream (garment plus target), applying separate AdaLN-zero normalization to each, and then running shared self-attention over all tokens, a single trainable network can preserve fine garment detail and temporal coherence. On image benchmarks the method reports the best SSIM, LPIPS, FID, and KID among compared methods, and on video benchmarks it improves SSIM, LPIPS, and VFID over dedicated video try-on systems. The authors interpret this as a \"renaissance\" of the single-network paradigm, with the dual-network approach as an unnecessary computational overhead.","pith_inferences":["The same \"normalize similar modalities together, separate dissimilar ones\" principle could apply to other multi-modal generation tasks, such as text-to-image editing or audio-visual generation, where a reference condition and a target share a modality.","The paper justifies the grouping with PCA visualizations and one ablation; an automated search over token-grouping strategies might find even better partitions when more modalities are added.","If the efficiency gain holds at production scale, the single-network paradigm could shift the cost model of e-commerce try-on services, since the ReferenceNet's memory and compute overhead is eliminated.","The claim that garment and target tokens should be normalized together is task-specific; for other tasks where the reference and target are not visually similar, the optimal grouping may differ."],"forward_implications":["A single architecture can now serve both image and video virtual try-on, with one model trained jointly on image and video data via position-embedding interpolation.","Without a ReferenceNet, trainable parameters drop substantially (e.g., 859.57M vs 1700.25M in the UNet comparison), so higher resolutions and longer clips become feasible.","The modality-split normalization rule transfers across backbone architectures: the same design works in a UNet and in a diffusion transformer, with the transformer version performing best.","Mask-free virtual try-on becomes practical: the trained model can generate paired video data that would otherwise be unavailable, enabling parsing-free video VTON."],"supporting_citations":[{"why":"CogVideoX serves as the pretrained diffusion transformer backbone that MN-VTON is initialized from and fine-tuned on.","marker":"[32]"},{"why":"VITONHD provides the high-resolution image VTON dataset used for training and evaluation.","marker":"[2]"},{"why":"DressCode supplies the second multi-category image dataset used in training and evaluation.","marker":"[16]"},{"why":"VIVID is the video VTON dataset and a dual-network baseline that this work compares against and jointly trains on.","marker":"[5]"},{"why":"VVT is the second video dataset and the source of the FW-GAN GAN-based baseline.","marker":"[25]"},{"why":"IDM-VTON is a dual-network image VTON baseline whose reported scores are used for comparison.","marker":"[3]"},{"why":"OOTDiffusion is a dual-network image VTON baseline included in quantitative and user-study comparisons.","marker":"[28]"},{"why":"CatVTON is a recent concatenation-based baseline that represents the competing single-network style approach.","marker":"[4]"}],"fun_headline_variants":["Single network rivals dual for try-on via modality normalization","No more ReferenceNet: single-network try-on with split norms","One network, two norms: try-on without ReferenceNet","Per-modality norms let single try-on network match dual systems"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole method hinges on the assumption that garment tokens and target person tokens belong to the same \"visual\" modality and should be normalized together, while text tokens should be normalized separately; the paper supports this with visualizations and a single ablation rather than a mechanism or exhaustive search.","fun_headline_variants_meta":{"raw":{"variants":["Single network rivals dual for try-on via modality normalization","No more ReferenceNet: single-network try-on with split norms","One network, two norms: try-on without ReferenceNet","Per-modality norms let single try-on network match dual systems"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000797,"raw_usage":{"total_tokens":3492,"prompt_tokens":916,"completion_tokens":2576,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":532,"completion_tokens_details":{"reasoning_tokens":2508}},"tokens_in":532,"tokens_out":2576,"duration_ms":18371,"temperature":1.0,"reasoning_tokens":2508,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:12:41.320579+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same MN-VTON pipeline with the alternative three-way split (text, garment, target each normalized independently) on VITONHD and VIVID test sets. If that configuration matches or beats the proposed two-way grouping on SSIM/LPIPS/VFID, the central claim that garment and target must share normalization collapses.","supporting_citations":[{"cited_title":"Viton-hd: High-resolution virtual try-on via misalignment-aware normalization","cited_arxiv_id":null,"evidence_quote":"VITONHD provides the high-resolution image VTON dataset used for training and evaluation."},{"cited_title":"Dress Code: High- Resolution Multi-Category Virtual Try-On","cited_arxiv_id":null,"evidence_quote":"DressCode supplies the second multi-category image dataset used in training and evaluation."},{"cited_title":"Fw-gan: Underwater image enhance- ment using generative adversarial network with multi-scale fusion","cited_arxiv_id":null,"evidence_quote":"VVT is the second video dataset and the source of the FW-GAN GAN-based baseline."},{"cited_title":"Catvton: Concatenation is all you need for virtual try-on with diffusion models, 2024","cited_arxiv_id":null,"evidence_quote":"CatVTON is a recent concatenation-based baseline that represents the competing single-network style approach."}],"review_version":1}