REVIEW 4 major objections 5 minor 47 references
Oxygen-TryOn claims to be the first any-item, multi-reference virtual try-on model to match or beat leading proprietary systems.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 07:03 UTC pith:HQGK7DLN
load-bearing objection Strong engineering and a genuinely different multi-reference conditioning scheme, but the headline SOTA claims lean on a self-judged Gemini protocol and an unreleased benchmark; worth refereeing, not worth believing as-is. the 4 major comments →
Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Oxygen-TryOn's central claim is that the right way to build a general try-on model is not to prompt a general image editor or inpaint inside a mask, but to train a fashion-native generator that reasons about what each reference item is, where it goes on the body, and how items layer and occlude. On the paper's own terms, the model—built on a multimodal LLM plus a multimodal diffusion transformer—accepts a variable number of references, encodes them into a shared token sequence, and synthesizes a dressed subject in a single pass, with the instruction binding each item to its position. The paper reports that on public benchmarks (DressCode, VITON-HD, TStars-VTON) and on its in-house Oxygen-Try
What carries the argument
The load-bearing design is the multi-reference conditioning scheme: subject and references are laid out as one token sequence, with the target region denoised from noise and reference regions kept as clean conditioning, separated only by intervals along the temporal axis of multimodal RoPE—no tags, segment embeddings, or masks. Item appearance travels in VAE latents; item identity and placement travel through the MLLM's semantic tokens and the natural-language instruction. On top of this, the training recipe is carried by a data engine that manufactures item–subject–result triplets, and by the RL stage's hybrid reward that fuses an in-house try-on reward model with a rubric-guided multimodal
Load-bearing premise
The load-bearing premise is that the rubric-guided judge used to produce the headline state-of-the-art scores is an honest external measure of try-on quality—but that same judge also supplies part of the RL reward during training; if the model has fitted the judge's rubric rather than human perception, the claimed margins over competitors would shrink or vanish under human evaluation, and the paper's own human study covers only 985 samples against two systems.
What would settle it
Run a pre-registered human preference study on the same 1,000 Oxygen-TryOn Bench samples plus a matched sample of TStars-VTON items, with annotators blind to model identity and with no overlap between raters and the reward-model training; if humans do not rank Oxygen-TryOn above GPT-Image-2 and Seedream5 Lite by margins comparable to the judged gap, the central SOTA claim is not supported.
If this is right
- Any wearable item—shoes, bags, hats, jewelry, as well as clothing—can be transferred onto a subject from either clean product shots or in-the-wild worn-on photos, in one pass.
- Multi-item composition lets a user assemble a full outfit (top, bottom, coat, hat, shoes, bag) while the model resolves layering and occlusion, not just swapping one garment.
- The same pass can follow general editing directives such as pose or background changes, so try-on and photo editing no longer require separate models.
- The documented data-engine and CPT–SFT–RL recipe gives the open community a path to reproduce or surpass closed proprietary try-on quality.
- The paper's own limitation section says five or more references still cause item confusion and dropped items, so the claimed capability is currently bounded at around four references.
Where Pith is reading between the lines
- Because the same rubric-guided judge is used both as an RL reward signal and as the scorer for the headline benchmark numbers, the reported margins may partly reflect alignment to that judge's rubric rather than to human perception; an independent, pre-registered human study with a judge not involved in training would settle whether real users see the same gap.
- If the mask-free, understanding-driven formulation is the real source of the gains, the recipe should transfer beyond fashion—for example, to product visualization, interior/decor transfer, and other object-on-subject generation tasks where identity and object fidelity must both be preserved.
- The paper's cross-domain demos (anime characters, oil paintings, film stills) suggest the model has learned item–subject wearing relations rather than a studio-photo prior; a systematic test on stylized domains would show how far that generalization extends.
- The stated bottleneck of five-plus references predicts a concrete scaling path: apply the same CPT–SFT–RL recipe to a base model with longer context or larger capacity, and evaluate whether multi-item consistency improves in step.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Oxygen-TryOn, a fashion-native foundation model for any-item virtual try-on. It reformulates try-on as a multi-reference, understanding-driven generation task rather than mask-based inpainting, building on JoyAI-Image-Edit with a Qwen3-VL MLLM, a Wan VAE, and a 16B MMDiT. Training uses a three-stage recipe: CPT, large-scale SFT, and RL with a hybrid reward combining an in-house try-on reward model and a rubric-guided Gemini 3.1 Pro judge. The authors report a data engine with over 50M raw images filtered to over 10M, multi-item pairing, and diverse category/subject coverage. Experiments cover DressCode/VITON-HD reconstruction metrics, TStars-VTON single- and multi-item VLM-judged scores, a private Oxygen-TryOn Bench, ablations of CPT/SFT/RL, and a human evaluation against two proprietary systems. The central claim is state-of-the-art consistency and realism for single-item try-on and leading multi-item composition, matching or surpassing proprietary systems and open-source models.
Significance. If the claims are substantiated, the contribution is significant: a unified any-item, multi-reference try-on system with a documented data engine and training recipe would be a useful advance over garment-centric inpainting methods. The paper has genuine strengths: the re-evaluation of baselines under a common protocol is transparent; the paired DressCode/VITON-HD results are strong; the architecture choice of separating reference semantics (MLLM) from item appearance (VAE latents) is well motivated; and the detailed description of the data engine and RL design is valuable for reproducibility. However, the headline single-item and multi-item SOTA numbers depend on a judge that is also used as an RL reward source, and the in-house benchmark is not released. These issues are load-bearing for the abstract-level claims, so the current evidence does not yet establish the stated SOTA.
major comments (4)
- [§3.2.3 and §4.1] Evaluator circularity: the rubric-guided Gemini 3.1 Pro judge used as one of the two RL reward sources in Section 3.2.3 is the same model used to score the TStars-VTON results in Tables 3–4 under the protocol described in Section 4.1. The policy is therefore directly optimized against the rubric of the judge that later produces the headline scores. This means the high TStars-VTON numbers may reflect reward overoptimization rather than human-perceived try-on quality. The human evaluation in Section 4.3 / Table 7 does not resolve the issue: it covers only 985 samples against two proprietary systems, reports no confidence intervals or significance tests, and the overall margin over GPT-Image-2 is 3.5502 vs 3.5375. Please provide independent human or judge-based evaluation on TStars-VTON (or a public subset), report uncertainty, and separate the evaluation judge from the RL reward model.
- [Table 3 and §4.2] The single-item SOTA claim is not established because TStars-Tryon1.0, the closest any-item baseline, has an officially reported Overall of 9.37, which exceeds Oxygen-TryOn's 9.36 in the same table. The paper lists this baseline under 'Officially Reported' and excludes it from the highlighted comparison because it cannot be reproduced. But the abstract and Section 1 claim state-of-the-art performance 'across public benchmarks.' Excluding the strongest official baseline changes the comparison set rather than establishing superiority. The claim should be reworded to 'best among re-evaluated baselines,' or the authors should obtain official-scoring results for TStars-Tryon1.0 and include them.
- [§4.1 and Table 5] Oxygen-TryOn Bench is an unreleased in-house benchmark, and its scores are produced by a GPT-5 judge with no external validation. The main deployment-oriented claims — usability rates of 86.79% and 85.43% versus 80.35% and 77.58% for GPT-Image-2 — rest entirely on this private benchmark. No error bars, confidence intervals, or significance tests are reported, and some differences are small (e.g., Cloth-to-Model Aesthetics: 3.736 vs. 3.768 for GPT-Image-2). To make the claimed real-world advantage verifiable, the benchmark, judge prompts, and scoring code should be released, and variance or significance information should be reported for at least the headline metrics.
- [§3.2.3] The hybrid reward is described only qualitatively as a 'weighted average' of the Gemini reward and the in-house try-on reward. The fusion weights, normalization, and group-relative reward details are not specified. Since Table 6 shows RL produces large gains (e.g., Model-to-Model usability 80.44→85.43), the relative contributions of the two reward sources are important for interpreting the improvement and for reproducing the recipe. Please report the fusion weights and, ideally, an ablation with each reward source removed.
minor comments (5)
- [Figure 7 caption vs §4.1] The caption says the TStars-VTON Overall score is 'aggregated by average,' while Section 4.1 states the Overall score is the geometric mean of the four dimensions. Please make these consistent.
- [Tables 2–3] The model name is misspelled as 'Oxygen-T ryOn' in Tables 2 and 3.
- [References] Reference [41] duplicates [40] (both are the Qwen-Image technical report), and reference [27] duplicates [3] (both FLUX.2). Please consolidate.
- [Table 7] The Overall score is described as a 'weighted average' but the weights are not reported. Please state the weights explicitly.
- [§6.3] The text says the model weights will be released, but the paper provides only a project page and no direct release link or license information. Please clarify the release status.
Circularity Check
Same Gemini 3.1 Pro judge is used as the RL reward source and as the TStars-VTON evaluator, so the headline single-item SOTA is partly forced by the training objective.
specific steps
-
fitted input called prediction
[Section 3.2.3 (Rubric-guided Gemini reward) and Section 4.1 (TStars-VTON evaluation)]
"In parallel, we query Gemini 3.1 Pro [19] as a multimodal judge for rubric-based RL. ... Following this protocol, we use Gemini 3.1 Pro as the judge."
The RL stage optimizes the policy with a hybrid reward whose components include the rubric-guided Gemini 3.1 Pro judge, fused with the in-house reward model via a weighted average. The same Gemini 3.1 Pro is then used as the judge for TStars-VTON, the benchmark that yields the paper's headline single-item SOTA claim (Overall 9.36 vs 8.77 next-best re-evaluated, Table 3). The evaluation instrument is thus a component of the training objective: high TStars-VTON scores are not an independent measure but reflect, at least in part, direct optimization against that judge's rubric. This is analogous to fitting a model to a target and then using that same target as the test prediction. The human study (985 samples, no significance tests) and the in-house GPT-5-judged bench provide partial independ
full rationale
The paper's central derivation is the CPT-SFT-RL recipe, and the RL stage clearly uses a hybrid reward whose two sources are the in-house reward model and Gemini 3.1 Pro as a rubric-guided judge. The evaluation section then adopts Gemini 3.1 Pro as the judge for TStars-VTON. Because the exact same judge is a training signal, the reported TStars-VTON superiority is, by construction, partly a measure of reward optimization rather than an independent external assessment. This is the paper's most load-bearing circular step: the abstract and Section 1 claim SOTA consistency/realism, and the TStars-VTON tables are the principal quantitative support for that claim. Other evidence reduces but does not eliminate the concern: DressCode/VITON-HD metrics are conventional paired/unpaired reconstruction scores; Oxygen-TryOn Bench uses GPT-5 as judge (not used in training); and a 985-sample human study against two proprietary systems is reported. These independent signals justify a partial, rather than total, circularity score. No additional self-citation chain appears load-bearing: JoyAI-Image-Edit [35] is cited as a base model, and the RL algorithms [17, 28, 46] are external methods, not uniqueness theorems imported to force the design. The exclusion of TStars-Tryon1.0 (not open-sourced; official Overall 9.37 exceeds Oxygen's 9.36) is a benchmark-completeness issue rather than a circularity issue. Overall, the derivation has real independent content, but the headline single-item SOTA is partially forced by the judge-in-training/evaluation overlap.
Axiom & Free-Parameter Ledger
free parameters (4)
- Reward fusion weights (Gemini vs in-house reward)
- Usability Rate threshold =
all three dimension scores strictly above 3
- CPT data mix ratio =
1:1 general-purpose:try-on
- In-house reward model head aggregation =
average of three reward heads
axioms (5)
- domain assumption JoyAI-Image-Edit provides a strong multimodal prior sufficient for try-on specialization
- domain assumption The synthetic data pipeline produces realistic supervision that transfers to real images
- domain assumption LLM judges (Gemini 3.1 Pro, GPT-5) yield valid, unbiased quality scores
- domain assumption Category/body-region compatibility rules capture valid wearing constraints
- standard math DiffusionNFT and Flow-GRPO provide stable RL optimization
invented entities (3)
-
Oxygen-TryOn Bench
no independent evidence
-
In-house try-on reward model
no independent evidence
-
Prompt Enhancer (PE)
no independent evidence
read the original abstract
We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for try-on through a dedicated data engine and try-on-specific training. Given one or more reference items (clean product shots or in-the-wild worn-on photos) and a single target subject image, it synthesizes a photorealistic image of the subject wearing the items across virtually any fashion category. Prior systems handle a single garment category in a studio setting, and recent multi-reference methods remain garment-centric; in contrast, Oxygen-TryOn supports diverse items and scenarios, including full- and half-body views, a variable number of references, and free multi-item composition, while faithfully preserving both subject identity and item appearance. Instead of mask-based inpainting, we reformulate try-on as a multi-reference, understanding-driven generation task. We build a data engine that collects, manufactures, annotates, and filters high-quality try-on data at scale, and design a three-stage recipe of continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage uses a hybrid reward combining an in-house try-on reward model with a proprietary, rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. It also follows general editing instructions (e.g., pose changes) in the same pass. Across public benchmarks and our in-house Oxygen-TryOn Bench, it achieves state-of-the-art consistency and realism on single-item try-on and leads on multi-item try-on, matching or surpassing both leading proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and open-source models (FLUX.2).
Reference graph
Works this paper leans on
-
[1]
Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025
Shuai Bai, Yuxuan Cai, Xionghui Chen, Qidong Huang, Kaixin Li, Zicheng Lin, Keming Zhu, et al. Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025. URLhttps://arxiv.org/abs/2511.21631
Pith/arXiv arXiv 2025
-
[2]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2. 5-vl technical report.arXiv:2502.13923, 2025
Pith/arXiv arXiv 2025
-
[3]
Flux.2: Towards interactive visual intelligence, 2025
Black Forest Labs. Flux.2: Towards interactive visual intelligence, 2025. URLhttps://bfl.ai/blog/flux2
2025
-
[4]
Deeper thinking, more accurate generation: Introducing seedream 5.0 lite, 2026
ByteDance. Deeper thinking, more accurate generation: Introducing seedream 5.0 lite, 2026. URL https://seed.bytedance.com/en/blog/
2026
-
[5]
Hunyuanimage 3.0 technical report.arXiv preprint arXiv:2509.23951, 2025
Siyu Cao, Hangting Chen, Peng Chen, et al. Hunyuanimage 3.0 technical report.arXiv preprint arXiv:2509.23951, 2025
Pith/arXiv arXiv 2025
-
[6]
Wear-any-way: Manipulable virtual try-on via sparse correspondence alignment
Mengting Chen, Xi Chen, Zhonghua Zhai, Chen Ju, Xuewen Hong, Jinsong Lan, and Shuai Xiao. Wear-any-way: Manipulable virtual try-on via sparse correspondence alignment. InEuropean Conference on Computer Vision, pages 124–142. Springer, 2024
2024
-
[7]
Tstars-tryon 1.0: Robust and realistic virtual try-on for diverse fashion items, 2026
Mengting Chen, Zhengrui Chen, Yongchao Du, Zuan Gao, Taihang Hu, Jinsong Lan, Chao Lin, Yefeng Shen, Xingjian Wang, Zhao Wang, Zhengtao Wu, Xiaoli Xu, Zhengze Xu, Hao Yan, Mingzhou Zhang, Jun Zheng, Qinye Zhou, Xiaoyong Zhu, and Bo Zheng. Tstars-tryon 1.0: Robust and realistic virtual try-on for diverse fashion items, 2026
2026
-
[8]
Viton-hd: High-resolution virtual try-on via misalignment-aware normalization
Seunghwan Choi, Sunghyun Park, Minsoo Lee, and Jaegul Choo. Viton-hd: High-resolution virtual try-on via misalignment-aware normalization. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14131–14140, 2021
2021
-
[9]
Improving diffusion models for authentic virtual try-on in the wild
Yisol Choi, Sangkyung Kwak, Kyungmin Lee, Hyungwon Choi, and Jinwoo Shin. Improving diffusion models for authentic virtual try-on in the wild. InEuropean Conference on Computer Vision (ECCV), 2024
2024
-
[10]
Catvton: Concatenation is all you need for virtual try-on with diffusion models, 2024
Zheng Chong, Xiao Dong, Haoxiang Li, Shiyue Zhang, Wenqing Zhang, Xujie Zhang, Hanqing Zhao, and Xiaodan Liang. Catvton: Concatenation is all you need for virtual try-on with diffusion models, 2024. URL https://arxiv.org/abs/2407.15886
Pith/arXiv arXiv 2024
-
[11]
Fastfit: Accelerating multi-reference virtual try-on via cacheable diffusion models,
Zheng Chong, Yanwei Lei, Shiyue Zhang, Zhuandi He, Zhen Wang, Xujie Zhang, Xiao Dong, Yiling Wu, Dongmei Jiang, and Xiaodan Liang. Fastfit: Accelerating multi-reference virtual try-on via cacheable diffusion models,
-
[12]
Paddleocr 3.0 technical report, 2025
Cheng Cui, Ting Sun, Manhui Lin, Tingquan Gao, Yubo Zhang, Jiaxuan Liu, Xueqing Wang, Zelun Zhang, Changda Zhou, Hongen Liu, Yue Zhang, Wenyu Lv, Kui Huang, Yichao Zhang, Jing Zhang, Jun Zhang, Yi Liu, Dianhai Yu, and Yanjun Ma. Paddleocr 3.0 technical report, 2025. URLhttps://arxiv.org/abs/2507.05595
Pith/arXiv arXiv 2025
-
[13]
Cheng Cui, Ting Sun, Suyin Liang, Tingquan Gao, Zelun Zhang, Jiaxuan Liu, Xueqing Wang, Changda Zhou, Hongen Liu, Manhui Lin, et al. Paddleocr-vl-1.5: Towards a multi-task 0.9 b vlm for robust in-the-wild document parsing.arXiv preprint arXiv:2601.21957, 2026
Pith/arXiv arXiv 2026
-
[14]
Aesthetic predictor v2.5.https://github.com/discus0434/aesthetic-predictor-v2-5, 2024
discus0434. Aesthetic predictor v2.5.https://github.com/discus0434/aesthetic-predictor-v2-5, 2024. SigLIP-based aesthetic score predictor. Accessed: 2026-06-30
2024
-
[15]
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. InICML, 2024
2024
-
[16]
Perceptual quality assessment of smartphone photography
Yuming Fang, Hanwei Zhu, Yan Zeng, Kede Ma, and Zhou Wang. Perceptual quality assessment of smartphone photography. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3677–3686, 2020
2020
-
[17]
Xiaolong Fu, Lichen Ma, Zipeng Guo, Gaojing Zhou, Chongxiao Wang, ShiPing Dong, Shizhe Zhou, Ximan Liu, Jingling Fu, Tan Lit Sin, et al. Dynamic-treerpo: Breaking the independent trajectory bottleneck with structured sampling.arXiv preprint arXiv:2509.23352, 2025
Pith/arXiv arXiv 2025
-
[18]
Introducing nano banana pro, 2025
Google. Introducing nano banana pro, 2025. URL https://blog.google/innovation-and-ai/products/nano-banana-pro/. Accessed: 2026-05-30. 29
2025
-
[19]
Gemini 3.1 pro: Announcing our latest gemini ai model
Google. Gemini 3.1 pro: Announcing our latest gemini ai model. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/, February
-
[20]
Any2anytryon: Leveraging adaptive position embeddings for versatile virtual clothing tasks
Hailong Guo, Bohan Zeng, Yiren Song, Wentao Zhang, Jiaming Liu, and Chuang Zhang. Any2anytryon: Leveraging adaptive position embeddings for versatile virtual clothing tasks. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 19085–19096, 2025
2025
-
[21]
Viton: An image-based virtual try-on network
Xintong Han, Zuxuan Wu, Zhe Wu, Ruichi Yu, and Larry S Davis. Viton: An image-based virtual try-on network. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 7543–7552, 2018
2018
-
[22]
Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020
2020
-
[23]
Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment.IEEE Transactions on Image Processing, 29:4041–4056, 2020
Vlad Hosu, Hanhe Lin, Tamas Sziranyi, and Dietmar Saupe. Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment.IEEE Transactions on Image Processing, 29:4041–4056, 2020
2020
-
[24]
Boyuan Jiang, Xiaobin Hu, Donghao Luo, Qingdong He, Chengming Xu, Jinlong Peng, Jiangning Zhang, Chengjie Wang, Yunsheng Wu, and Yanwei Fu. Fitdit: Advancing the authentic garment details for high-fidelity virtual try-on.arXiv preprint arXiv:2411.10499, 2024
Pith/arXiv arXiv 2024
-
[25]
Apddv2: Aesthetics of paintings and drawings dataset with artist labeled scores and comments.Advances in Neural Information Processing Systems, 37:103064–103075, 2024
Xin Jin, Qianqian Qiao, Yi Lu, Huaye Wang, Heng Huang, Shan Gao, Jianfei Liu, and Rui Li. Apddv2: Aesthetics of paintings and drawings dataset with artist labeled scores and comments.Advances in Neural Information Processing Systems, 37:103064–103075, 2024
2024
-
[26]
Jeongho Kim, Hoiyeong Jin, Sunghyun Park, and Jaegul Choo. Promptdresser: Improving the quality and controllability of virtual try-on via generative textual prompt and prompt-aware mask.arXiv preprint arXiv:2412.16978, 2024
Pith/arXiv arXiv 2024
-
[27]
FLUX.2: State-of-the-Art Visual Intelligence.https://bfl.ai/blog/flux-2, 2025
Black Forest Labs. FLUX.2: State-of-the-Art Visual Intelligence.https://bfl.ai/blog/flux-2, 2025
2025
-
[28]
Flow-grpo: Training flow matching models via online rl.arXiv preprint arXiv:2505.05470, 2025
Jie Liu, Gongye Liu, Jiajun Liang, Yangguang Li, Jiaheng Liu, Xintao Wang, Pengfei Wan, Di Zhang, and Wanli Ouyang. Flow-grpo: Training flow matching models via online rl.arXiv preprint arXiv:2505.05470, 2025
Pith/arXiv arXiv 2025
-
[29]
Hpsv3: Towards wide-spectrum human preference score
Yuhang Ma, Xiaoshi Wu, Keqiang Sun, and Hongsheng Li. Hpsv3: Towards wide-spectrum human preference score. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 15086–15095, 2025
2025
-
[30]
Dress code: High-resolution multi-category virtual try-on
Davide Morelli, Matteo Fincato, Marcella Cornia, Federico Landi, Fabio Cesari, and Rita Cucchiara. Dress code: High-resolution multi-category virtual try-on. InComputer Vision – ECCV 2022, pages 345–362, 2022
2022
-
[31]
Gpt-image-1.5 model card, 2025
OpenAI. Gpt-image-1.5 model card, 2025. URLhttps://platform.openai.com/docs/models/gpt-image-1-5
2025
-
[32]
Gpt-image-2 model card, 2026
OpenAI. Gpt-image-2 model card, 2026. URLhttps://platform.openai.com/docs/models/gpt-image-2
2026
-
[33]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022
2022
-
[34]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024
Pith/arXiv arXiv 2024
-
[35]
Lin Song, Wenbo Li, Guoqing Ma, Wei Tang, Bo Wang, Yuan Zhang, Yijun Yang, Yicheng Xiao, Jianhui Liu, Yanbing Zhang, et al. Joyai-image: Awaking spatial intelligence in unified multimodal understanding and generation.arXiv preprint arXiv:2605.04128, 2026
Pith/arXiv arXiv 2026
-
[36]
Firered-image-edit-1.0 techinical report.arXiv preprint arXiv:2602.13344, 2026
Super Intelligence Team, Changhao Qiao, Chao Hui, Chen Li, Cunzheng Wang, Dejia Song, Jiale Zhang, Jing Li, Qiang Xiang, Runqi Wang, et al. Firered-image-edit-1.0 techinical report.arXiv preprint arXiv:2602.13344, 2026
arXiv 2026
-
[37]
Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Pith/arXiv arXiv 2025
-
[38]
Toward characteristic-preserving image-based virtual try-on network
Bochao Wang, Huabin Zheng, Xiaodan Liang, Yimin Chen, Liang Lin, and Meng Yang. Toward characteristic-preserving image-based virtual try-on network. InProceedings of the European conference on computer vision (ECCV), pages 589–604, 2018. 30
2018
-
[39]
Retouchformer: semi-supervised high-quality face retouching transformer with prior-based selective self-attention
Xue Wen, Lianxin Xie, Le Jiang, Tianyi Chen, Si Wu, Cheng Liu, and Hau-San Wong. Retouchformer: semi-supervised high-quality face retouching transformer with prior-based selective self-attention. InProceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 5903–5911, 2024
2024
-
[41]
Qwen-image technical report.arXiv preprint arXiv:2508.02324, 2025
Chenfei Wu, Jiahao Li, Jingren Zhou, Junyang Lin, et al. Qwen-image technical report.arXiv preprint arXiv:2508.02324, 2025
Pith/arXiv arXiv 2025
-
[42]
Keming Wu, Sicong Jiang, Max Ku, Ping Nie, Minghao Liu, and Wenhu Chen. Editreward: A human-aligned reward model for instruction-guided image editing.arXiv preprint arXiv:2509.26346, 2025
arXiv 2025
-
[43]
Ootdiffusion: Outfitting fusion based latent diffusion for controllable virtual try-on
Yuhao Xu, Tao Gu, Weifeng Chen, and Arlene Chen. Ootdiffusion: Outfitting fusion based latent diffusion for controllable virtual try-on. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 8996–9004, 2025
2025
-
[44]
Tunnel try-on: Excavating spatial-temporal tunnels for high-quality virtual try-on in videos
Zhengze Xu, Mengting Chen, Zhao Wang, Linyu Xing, Zhonghua Zhai, Nong Sang, Jinsong Lan, Shuai Xiao, and Changxin Gao. Tunnel try-on: Excavating spatial-temporal tunnels for high-quality virtual try-on in videos. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 3199–3208, 2024
2024
-
[45]
Retouchgpt: Llm-based interactive high-fidelity face retouching via imperfection prompting
Wen Xue, Chun Ding, Ruotao Xu, Si Wu, Yong Xu, and Hau-San Wong. Retouchgpt: Llm-based interactive high-fidelity face retouching via imperfection prompting. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 9059–9067, 2025
2025
-
[46]
Kaiwen Zheng, Huayu Chen, Haotian Ye, Haoxiang Wang, Qinsheng Zhang, Kai Jiang, Hang Su, Stefano Ermon, Jun Zhu, and Ming-Yu Liu. Diffusionnft: Online diffusion reinforcement with forward process.arXiv preprint arXiv:2509.16117, 2025
Pith/arXiv arXiv 2025
-
[47]
Zijian Zhou, Shikun Liu, Xiao Han, Haozhe Liu, Kam Woh Ng, Tian Xie, Yuren Cong, Hang Li, Mengmeng Xu, Juan-Manuel Pérez-Rúa, Aditya Patel, Tao Xiang, Miaojing Shi, and Sen He. Learning flow fields in attention for controllable person image generation.arXiv preprint arXiv:2412.08486, 2024. 31
Pith/arXiv arXiv 2024
-
[2025]
URLhttps://arxiv.org/abs/2508.20586
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.