{"id":"56e33b74-77b7-4f0b-887b-5fc23a7002a4","arxiv_id":"2605.12939","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"DirectTryOn achieves state-of-the-art one-step virtual try-on performance by applying pure conditional transport, garment preservation loss, and self-consistency loss to straighten trajectories in pretrained generative models.","lead":"DirectTryOn proposes a one-step virtual try-on method that straightens conditional sampling trajectories in diffusion models using targeted losses and distillation. This targets the high inference cost of multi-step VTON while maintaining quality for constrained garment placement tasks.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Key claim that deviation from straight paths stems mainly from pretrained-model mismatch (not task itself) lacks isolation via from-scratch baseline.","rationale":"The reader's weakest assumption directly matches the paper's motivating insight about conditional straightness. The proposed check isolates whether that straightness is an intrinsic property of the task or an artifact of the fine-tuning procedure, which is the precise point where the central claim could fail while still producing strong empirical numbers.","tokens_in":1702,"tokens_out":416,"duration_ms":33047,"concrete_test":"On a 10k-image subset of the VTON training data, train a small conditional U-Net from random initialization using only the standard flow-matching or diffusion objective (no proposed losses or distillation). After convergence, measure (a) average trajectory curvature (integral of ||v_t - v_{t+dt}|| normalized by path length) and (b) FID/LPIPS of one-step vs. 10-step sampling on the test set. If one-step quality remains substantially worse than the fine-tuned model and curvature stays high, the mismatch hypothesis is supported; if one-step quality approaches the reported SOTA, the claim requires revision.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central argument rests on the observation that VTON's heavy conditioning makes trajectories inherently straighter than unconditional generation, so the main obstacle is the pretrained base model's objective rather than the try-on task. The three modifications (pure conditional transport, garment preservation loss, self-consistency loss) plus distillation are presented as correcting this mismatch. However, because the paper states that limited task-specific data makes training from scratch impractical, no direct comparison exists that holds architecture and data fixed while removing the pretrained initialization. Without that control, improvements could arise from the auxiliary losses regularizing the fine-tuning rather than revealing or exploiting an intrinsically straighter conditional manifold. This leaves the load-bearing premise—that one-step sampling is a natural solution once the mismatch is removed—empirically unseparated from the specific training recipe.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces DirectTryOn for one-step virtual try-on (VTON) by straightening conditional transport trajectories in pretrained diffusion/flow models. It claims that VTON's heavy conditioning makes trajectories inherently straighter than in unconditional generation, so the main obstacle is pretrained-model mismatch rather than the task; three modifications (pure conditional transport, garment preservation loss, self-consistency loss) plus one-step distillation are proposed to correct this and achieve SOTA one-step performance.","tokens_in":1862,"tokens_out":350,"duration_ms":52239,"significance":"If the central claim holds, the work would be significant for efficient VTON by exploiting task-specific trajectory properties to reduce inference from multi-step to single-step sampling while preserving quality, with practical value for real-time applications such as e-commerce.","major_comments":[{"comment":"Abstract: the load-bearing premise that 'the deviation from an ideal straight path mainly comes from the mismatch between pretrained base models and the conditional nature of try-on generation, rather than from the task itself' is not isolated, because no from-scratch baseline (holding architecture and data fixed) is reported despite the acknowledgment that limited task-specific data makes such training impractical; without this control, observed gains cannot be attributed specifically to revealing an intrinsically straighter conditional manifold versus the regularizing effect of the auxiliary losses.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract and introduction would benefit from explicit quantitative statements of the step reduction (e.g., from N to 1) and the exact metrics where SOTA is claimed.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. We address the major comment below.","responses":[{"response":"We agree that a from-scratch baseline holding architecture and data fixed would provide the cleanest isolation of whether the conditional manifold is intrinsically straighter. As the manuscript already states, however, the scarcity of high-quality paired garment-person data renders training from scratch impractical both in data volume and compute. This constraint is why virtually all recent VTON methods, including strong baselines, start from the same class of pretrained models. Our ablations (Section 4.3) isolate the contribution of each component: ablating pure conditional transport, garment preservation loss, or self-consistency loss individually increases trajectory curvature and degrades one-step FID/LPIPS, while the full combination yields the reported gains. These components are not generic regularizers; they explicitly target the pretrained-conditional mismatch. We also outperform other methods that fine-tune the identical pretrained backbones without straightening. In revision we will expand the abstract and Section 3 to explicitly discuss this limitation and the supporting ablation evidence.","revision_made":"partial","referee_comment":"[Abstract] Abstract: the load-bearing premise that 'the deviation from an ideal straight path mainly comes from the mismatch between pretrained base models and the conditional nature of try-on generation, rather than from the task itself' is not isolated, because no from-scratch baseline (holding architecture and data fixed) is reported despite the acknowledgment that limited task-specific data makes such training impractical; without this control, observed gains cannot be attributed specifically to revealing an intrinsically straighter conditional manifold versus the regularizing effect of the auxiliary losses."}],"tokens_in":1289,"tokens_out":382,"duration_ms":38751,"standing_objections":["A from-scratch baseline is not feasible due to limited task-specific paired data, as already noted in the manuscript."]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that this work gets solid one-step results on virtual try-on by treating the task as heavily conditioned generation and adding three changes: pure conditional transport, a garment preservation loss, and a self-consistency loss, plus a distillation stage. That combination looks new for this specific setting and directly targets inference cost, which matters for real e-commerce use. The abstract makes a clear case that VTON constraints should allow straighter paths than general image generation, and the experiments are presented as confirming SOTA one-step quality. That part is useful and worth checking against existing multi-step and accelerated baselines. The soft spot is the one the stress-test flags. Without a from-scratch run (which the authors say is impractical due to data limits), it is hard to separate whether the gains come from exploiting an inherently straighter conditional manifold or from the new losses simply regularizing fine-tuning better. The paper acknowledges the data constraint, so the gap is understandable, but it leaves the load-bearing premise a bit indirect. I would also like to see explicit trajectory metrics, such as average path curvature or step-wise deviation, rather than relying only on final image scores. This is for people working on efficient conditional generation and fashion applications who need faster sampling without big quality drops. The ideas are grounded enough in the diffusion and flow literature to merit referee time, even if the isolation of the central mechanism needs tightening in revision. I would send it to review rather than desk reject.","headline":"The paper shows a workable path to one-step VTON by adding targeted losses and distillation to straighten conditional trajectories, but the claim that this straightness is mainly a pretrained-model mismatch lacks a clean control experiment.","tokens_in":2359,"tokens_out":371,"would_cite":false,"duration_ms":36305,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"One-step VTON flow-matching paper has no structural overlap with RS distinction-to-physics chain","alignment":"orthogonal","rationale":"Central machinery (pure conditional transport, garment preservation loss, self-consistency loss, LADD distillation on MMDiT) operates entirely within rectified-flow / diffusion acceleration literature; no J-cost, phi-ladder, 8-tick periodicity, ratio-symmetric forcing, or any theorem from the RS modules (e.g., Cost.FunctionalEquation, Foundation.RealityFromDistinction, AlexanderDuality) appears or is paralleled. Domain (cs.CV image synthesis) is outside RS scope.","tokens_in":52064,"confidence":"high","tokens_out":149,"duration_ms":8749,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Virtual try-on can reach state-of-the-art quality in one sampling step by straightening the conditional transport path.","keywords":["virtual try-on","one-step sampling","conditional transport","diffusion models","image generation","garment preservation","efficient inference"],"falsifier":"A direct comparison showing that the one-step outputs are visibly inferior to multi-step outputs from the same model in terms of garment alignment or realism would falsify the claim.","tokens_in":2591,"feed_emoji":"👕","tokens_out":484,"duration_ms":45508,"temperature":0.7,"pith_summary":"The paper shows that virtual try-on generation differs from general image synthesis because the output is tightly constrained by the input person and garment images. This constraint allows the sampling trajectory to be made much straighter than usual. By introducing pure conditional transport, a garment preservation loss, and a self-consistency loss, followed by one-step distillation, the method trains a model that produces high-quality try-on results directly in a single step. This avoids the high cost of multi-step sampling in existing diffusion and flow-based approaches while matching or exceeding their performance.","feed_headline":"One-step virtual try-on matches multi-step results","feed_subtitle":"Straightening the conditional path with targeted losses allows high-quality try-on images in a single diffusion step.","key_machinery":"Straightened conditional transport achieved through pure conditional transport, garment preservation loss, self-consistency loss, and one-step distillation.","core_discovery":"The central discovery is that the deviation from straight paths in try-on comes from the mismatch with pretrained models rather than the task itself, so targeted modifications—pure conditional transport, garment preservation loss, and self-consistency loss—combined with one-step distillation enable accurate one-step virtual try-on.","pith_inferences":["Similar trajectory straightening may apply to other image-to-image translation tasks with strong conditional constraints.","Future work could explore whether this approach reduces the need for large pretrained models in specific domains.","Testing on diverse body types and garment styles would reveal the limits of the straight-path assumption."],"forward_implications":["High-quality virtual try-on becomes feasible at real-time speeds.","Existing pretrained generative models can be adapted for efficient conditional tasks without full retraining.","Sampling efficiency improves without sacrificing output fidelity in constrained generation settings.","Virtual try-on systems can be deployed on devices with limited compute."],"fun_headline_variants":["Straightened paths enable one-step virtual try-on","Targeted losses straighten conditional VTON trajectories","Pure transport yields single-step sampling accuracy","Model mismatch fixed via consistency and garment losses"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The outputs in virtual try-on are sufficiently constrained by the input conditions that a straight sampling path suffices for high quality.","fun_headline_variants_meta":{"raw":{"variants":["Straightened paths enable one-step virtual try-on","Targeted losses straighten conditional VTON trajectories","Pure transport yields single-step sampling accuracy","Model mismatch fixed via consistency and garment losses"]},"model":"grok-4.3","cost_usd":0.004302,"raw_usage":{"total_tokens":2062,"prompt_tokens":629,"num_sources_used":0,"completion_tokens":53,"cost_in_usd_ticks":43015500,"prompt_tokens_details":{"text_tokens":629,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1380,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":629,"tokens_out":53,"duration_ms":38141,"temperature":1.0,"reasoning_tokens":1380,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-14T19:56:43.605891+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A direct comparison showing that the one-step outputs are visibly inferior to multi-step outputs from the same model in terms of garment alignment or realism would falsify the claim.","supporting_citations":[],"review_version":1}