Pith. sign in

REVIEW 4 major objections 5 minor 153 references

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget

T0 review · 4 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read Boogu-Image-0.1 argues that upgrading the understanding side of a text-to-image system—encoder, prompt rewriting, captions, routing—can make a model trained on 208.62M images at roughly $400K competitive with far costlier open- and closed-s

desk verdict Genuinely useful engineering report with honest ablations; flagship SOTA claim rests on an in-house benchmark needing external validation. read the letter →

arxiv 2607.13125 v2 pith:2JDMA2ZH submitted 2026-07-14 cs.CV cs.AI

classification cs.CVcs.AI
keywords text-to-imagegenerationimageeditingmultimodalunderstandingagenticpromptrewritingtrainingefficiencydatacurationEloevaluationbilingualtextrendering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper's central claim is that image-generation quality is gated less by generator size or data volume than by how well the system understands instructions. It builds a unified understanding-and-generation family around a stronger instruction encoder, an agentic prompt rewriter, captions designed around user demand, and a model router, and reports that this lets a from-scratch 10B-parameter model trained on only 208.62M unique images match or beat other open models and approach closed frontier systems. On its own in-house blind-vote benchmark and on a post-freeze open benchmark, the models lead the open-source tier, with the thinking variant adding the largest gains in text rendering and creativity. The authors also report practical data findings: a curated 47M-image syllabus beats 187M raw open-source images, roughly 300 exposures reliably teach a Chinese glyph, and capping the timestep-shift at the 1K token-count regime fixes slow convergence at 2K. The authors explicitly caution that their primary arena is in-house and that the editing benchmark inverts under human evaluation, so the headline rankings rest on instruments that await independent validation.

What carries the argument

The load-bearing mechanism is a decoupled understanding stack placed upstream of generation. A mid-sized frozen vision-language model serves as the instruction encoder, treated as a sensor whose capacity upper-bounds what the diffusion transformer can receive, while a separate reasoning-capable agentic rewriter and a per-aspect captioning pipeline supply the conditioning and supervision. Completing the stack is a router that classifies request difficulty and selects fast versus heavy model variants, turning inference time into a dial on a quality-latency curve. Two supporting techniques carry much of the training-efficiency claim: a syllabus-guided data curriculum with explicit defect captio

What would settle it

Run a public, independent blind pairwise comparison, for example on a held-out bilingual prompt set, between Boogu-Image-0.1-Turbo-Thinking and the leading open baselines it claims to beat, and check whether its Elo still leads; since the weights are released, this is directly executable.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that understanding can be engineered as a first-class, decoupled component of a generation system rather than left implicit in a large generator. It decomposes understanding into an instruction encoder that acts as a lossy sensor, an agentic prompt rewriter that translates intent without embellishing clear prompts, captions that name what users will ask for (including visual flaws), and a difficulty-aware router that dispatches easy requests to fast models. The empirical core is a series of controlled comparisons: encoder scale alone moves GenEval from 0.60 to 0.65 with no saturation; the same 47M-image syllabus outperforms 187M of open-source data; th

Load-bearing premise

The headline top-open-source-model ranking depends on Boogu Arena, an in-house Elo benchmark whose agreement with public human preference was measured on only five models; if that arena does not reflect real user preference, the central competitive claim is unsupported.

Editorial extensions

If this is right

  • Open-source teams can budget for roughly 200M unique images and about $400K of compute and still reach the top of the open tier, provided understanding-side components are treated as part of the model.
  • Evaluation that reports quality without latency is incomplete: the same Boogu model occupies many points on the quality-time curve, and thinking and routing variants make that curve explicit.
  • Capabilities that are not named in captions cannot be recovered at inference time, so captioning should be planned per anticipated user demand rather than delegated to a single vision-language model.
  • Saturated academic benchmarks such as GenEval and DPG-Bench no longer track human preference, so the paper's results on its arena and on a post-freeze open benchmark carry the evidential weight.
  • Skills embedded in the rewriter, such as counting, infographic layout, reasoning, NSFW handling, and scene text, transfer into measurable generation gains in those same dimensions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the roughly 300-exposure glyph result generalizes, it predicts a cheap and testable recipe for any new writing system or token set: enumerate the vocabulary and floor each unit's exposure rather than scaling raw data.
  • The reported inversion on the editing benchmark, where a closed system scores lower yet wins in the authors' own human evaluation, suggests that model-based editing metrics systematically compress quality gaps; a public human editing arena could settle which ranking is real.
  • The finding that heavy aesthetic reinforcement learning narrows output distributions implies a policy tension: optimizing average appeal can erase off-mainstream demographics and styles, which matters for fairness and access as much as for diversity.
  • Releasing weights with full recipes makes the understanding-first budget claim directly falsifiable and re-runnable, which is the main reason the cost figure can be checked rather than taken on faith.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces Boogu-Image-0.1, an open-source unified text-to-image and image-editing model family (Base, Turbo, Edit, Edit-Turbo), and argues that strengthening the understanding component — via a stronger frozen multimodal encoder (Qwen3-VL-8B), an agentic prompt rewriter, a curated data syllabus, and agentic inference-time scaling — substantially improves generation quality under a minimal compute budget (208.62M unique images, ~$400K theoretical training cost). The evaluation relies on an in-house blind pairwise benchmark (Boogu Arena), on external benchmarks Qwen-Image-Bench and LongText-Bench, and on ImgEdit-Bench for editing, claiming top-tier open-source status and performance approaching closed-source systems.

Significance. If the central claims hold, the paper offers two useful contributions: (i) a systematic demonstration, with ablations, that instruction-encoder scale, curated data syllabus, and agentic prompt rewriting each yield measurable gains under a fixed budget, and (ii) a practical open release of weights, code, and training recipes. The ablations in Tables 7–8 and Figure 18 are genuine and non-circular, comparing actual checkpoints and data configurations. The bottleneck is the headline positioning: the claim of "approaching closed-source systems" and of open-source leadership on human preference rests almost entirely on an in-house benchmark whose external validation is thin, and the paper itself documents an inversion between its automated editing benchmark and its own human evaluation. The external Qwen-Image-Bench results support open-source leadership only among the subset of models evaluated and do not support the closed-source gap claim.

major comments (4)
  1. [§2.2.1, Figure 7] The Boogu Arena Elo ranking is the primary evidence for the headline "approaching closed-source" claim, but its external validation is limited to five models, none of which are Boogu models. With n=5, a Pearson r of 0.986 and Spearman 1.0 are consistent with a linear trend but do not establish that Boogu's absolute Elo position transfers to the LMArena scale; the Boogu points are extrapolations. The paper reports no confidence intervals on Elo, no inter-annotator agreement, and no independent annotation protocol, and the prompts are only promised for future release. The paper itself (Sec. 2.3.1) shows that its in-house preference instrument can invert the ranking of two models relative to human evaluation, so Boogu Arena cannot currently carry the closed-source gap claim.
  2. [§2.3.1, Table 6] The text states that Boogu-Image-0.1-Edit-Thinking "attains the best overall score" on ImgEdit-Bench, but the immediately following Discussion and Limitations says that in the authors' own human evaluations Nano-Banana-Pro outperforms Boogu despite scoring lower (4.37 vs 4.64), and the table is labeled "for reference only." This is not a minor caveat: the abstract and introduction use "matches or surpasses" across standard benchmarks without this qualification, and the internal inversion means the ImgEdit ranking should not be presented as a headline result. The claim needs to be recast or the human evaluation data reported.
  3. [§2.2.2, Tables 1–2; §2.2.3, Table 3] The open-source leadership claim on Qwen-Image-Bench and LongText-Bench is limited by the evaluated set: the text says "due to time constraints, the evaluation does not yet cover all available open-source baselines," and no inference-time budget is reported for any model. Since the paper argues (Sec. 3.2.1) that quality must be reported jointly with latency, the thinking variants, which use agentic rewriting and reflection, should be compared with closed-source systems under a matched or explicitly reported inference budget; otherwise the 'approaching closed-source' portion of the claim is not supported by these tables.
  4. [§3.2.3 (Dynamic time shifting) and Figure 32] The rectified dynamic time-shifting scheme introduces a free parameter ncap (e.g., 4096), and its benefit is illustrated by quantile plots rather than by a controlled training ablation showing that the rectified scheme improves final generation metrics (e.g., FID or benchmark scores) relative to standard dynamic shifting at 2K. The claim that this is a decisive detail for the Boogu pipeline is plausible but currently under-supported; either add an ablation or soften the presentation.
minor comments (5)
  1. [Figure 1] The right panel lists 'Boogu-Image-0.1-Turbo-PE' but the model family in the abstract and Section 2 includes Base, Turbo, Edit, Edit-Turbo; PE is not defined in the caption or text. Please clarify or rename.
  2. [Tables 1–6] Model names are inconsistent: 'Nano-Banana-Pro' vs 'NanoBanana-Pro', and 'Qwen-Image-2.0-2026-03-03' vs 'Qwen-Image-2.0'. Also, some tables say 'Our model series achieves...' while the model is highlighted in red; ensure the caption convention is uniform.
  3. [§2.2.4] The decision to report GenEval and DPG-Bench 'for reference only' is reasonable, but the same standards should apply to ImgEdit-Bench, which is also based on VLM scoring and acknowledged to invert human preference. Consider adding a unified statement about which benchmarks the authors regard as primary and why.
  4. [§4] The limitations paragraph is candid but does not mention the absence of an external human-preference validation for Boogu Arena; this is the most load-bearing limitation for the abstract's claim and should be stated there.
  5. [Abstract] The phrase 'the base model's theoretical training cost is only approximately $400K' is used without defining what 'theoretical' includes (hardware, electricity, ablations, data curation). Please provide a cost breakdown or soften the claim.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: central claims rest on independent ablations and external benchmarks; Boogu Arena is a validity concern, not a circular reduction.

full rationale

The paper's load-bearing derivations are direct empirical ablations over actual checkpoints and data configurations: Table 7 varies the frozen instruction encoder and measures GenEval; Figure 18 scales the rewriter backbone and reports a defined advantage score; Table 8 compares training on 187M open-source data vs. the Boogu Syllabus on Qwen-Image-Bench and LongText-Bench; Figures 27-28 and Table 9 measure glyph-exposure thresholds. In each case the manipulated variable is not defined in terms of the measured outcome, so no self-definitional reduction is present. The supporting benchmarks (Qwen-Image-Bench, LongText-Bench, GenEval, DPG-Bench) are external to the paper. The Boogu Arena leaderboard is built in-house, but its Elo scores are not fitted to Boogu outputs, and the LMArena agreement check (Figure 7), while based on only five non-Boogu models, is a calibration exercise rather than a definitional reduction. The only apparent self-citation ([78], used as an example in the open-challenges discussion) is not load-bearing. Benchmark-validity concerns about Boogu Arena are legitimate evaluation risks, but they are not circularity under the definitions used here.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central claim rests on a few domain assumptions about evaluation validity and generality of exposure thresholds; no free parameters are fitted to make the derivation work, and no new physical entities are introduced.

free parameters (1)
  • ncap (dynamic time-shift clamp) = 4096
    Chosen by hand to inherit 1K-resolution training behavior for 2K training; not part of the claim-derivation, but an ad hoc training choice reported in Section 3.2.3.
assumptions (4)
  • domain assumption The Boogu Arena Elo computed from in-house blind pairwise votes is a faithful proxy for real-world human preference (correlation with LMArena r=0.986 on five anchor models).
    Used to rank Boogu models as top open-source in Section 2.2.1/Figure 6. No inter-annotator agreement or external annotator independence is reported.
  • domain assumption VLM/automated benchmarks (Qwen-Image-Bench, LongText-Bench, ImgEdit-Bench) are meaningful for the claims despite acknowledged saturation and their own ImgEdit inversion.
    Tables 1-6 use these scores as headline evidence; Section 2.3.1 concedes the ImgEdit ranking contradicts human preference.
  • domain assumption A per-exposure threshold (300 per glyph; ~4,500 per identity) measured on a small subset generalizes to the full training distribution.
    Section 3.2.2-3.2.3 derive quantitative claims from ~17 Chinese characters and one volunteer identity.
  • standard math Standard flow-matching/diffusion and DiT formulations are correct as used.
    The training and BOG sections rely on established diffusion/flow-matching math (Sections B and C).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget." pith.science (2026). https://pith.science/paper/2JDMA2ZH

@misc{pith2026260713125,
  author       = {Pith},
  title        = {Pith review of: Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2JDMA2ZH}},
  note         = {Machine review of arXiv:2607.13125}
}
abstract

We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual (Chinese-English) text rendering. Closed-source multimodal systems like Nano-Banana-Pro and GPT-Image-2 achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. In this work, we demonstrate that strengthening the understanding capability of the system, through a stronger multimodal encoder, agentic prompt rewriting, and related techniques, together with improvements in data quality, training pipelines, and agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets. Comprehensive evaluations show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks, and achieves results approaching leading closed-source systems. Notably, this is accomplished with only 208.62 million unique images. The base model's theoretical training cost is only approximately \$400K. We share practical discussions that we believe are valuable to the broader research community, and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. Our code is available here: https://github.com/Boogu-Project/Boogu-Image.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

153 extracted references · 60 linked inside Pith

  1. [1]

    Arena ai: The official ai ranking & llm leaderboard.https://arena.ai/, 2026

  2. [2]

    Ideogram 4.https://ideogram.ai/blog/ideogram-4.0/, 2026

    Ideogram AI. Ideogram 4.https://ideogram.ai/blog/ideogram-4.0/, 2026

  3. [3]

    Qwen3-vl technical report, 2025

    Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan ...

  4. [4]

    Qwen2.5-vl technical report, 2025

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report, 2...

  5. [5]

    Beyond the imitation game: Quantifying and extrapolating the capabilities of language models

    BIG bench authors. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URLhttps://openreview.net/ forum?id=uyTL5Bvosj

  6. [6]

    Improving image generation with better captions.Computer Science

    James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al. Improving image generation with better captions.Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3):8, 2023

  7. [7]

    Announcing Black Forest Labs: Black forest labs’ inaugural model suite, FLUX.1, 2024

    Black Forest Labs. Announcing Black Forest Labs: Black forest labs’ inaugural model suite, FLUX.1, 2024. URL https://blackforestlabs.ai/announcing-black-forest-labs/. Accessed: 2026-04-25

  8. [8]

    FLUX.2: Analyzing and enhancing the latent space of FLUX – representation comparison, 11

    Black Forest Labs. FLUX.2: Analyzing and enhancing the latent space of FLUX – representation comparison, 11

Show all 153 references
  1. [9]

    Flux.2: Production-grade ai image generation and editing model with 4mp photorealistic output and multi-reference control.https://bfl.ai/models/flux-2, 2025

    Black Forest Labs. Flux.2: Production-grade ai image generation and editing model with 4mp photorealistic output and multi-reference control.https://bfl.ai/models/flux-2, 2025

  2. [10]

    Large scale GAN training for high fidelity natural image synthesis

    Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale GAN training for high fidelity natural image synthesis. In International Conference on Learning Representations, 2019. URL https://openreview.net/ forum?id=B1xsqj09Fm

  3. [11]

    Instructpix2pix: Learning to follow image editing instructions

    Tim Brooks, Aleksander Holynski, and Alexei A Efros. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 18392–18402, 2023

  4. [12]

    Video generation models as world simulators.OpenAI Blog, 1(8):1, 2024

    Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Leo Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, et al. Video generation models as world simulators.OpenAI Blog, 1(8):1, 2024

  5. [13]

    Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset, 2022

    Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset, 2022

  6. [14]

    Seedream 4.5.https://seed.bytedance.com/en/seedream4_5, 2026

    ByteDance Seed Team. Seedream 4.5.https://seed.bytedance.com/en/seedream4_5, 2026

  7. [15]

    Seedream 5.0 lite

    ByteDance Seed Team. Seedream 5.0 lite. https://seed.bytedance.com/en/blog/ deeper-thinking-more-accurate-generation-introducing-seedream-5-0-lite, 2026

  8. [16]

    Mm-iq: Benchmarking human-like abstraction and reasoning in multimodal models

    Huanqia Cai, Yijun Yang, and Winston Hu. Mm-iq: Benchmarking human-like abstraction and reasoning in multimodal models. arXiv preprint arXiv:2502.00698, 2025

  9. [17]

    Hidream-o1-image: A natively unified image generative foundation model with pixel-level unified transformer.arXiv preprint arXiv:2605.11061, 2026

    Qi Cai, Jingwen Chen, Chengmin Gao, Zijian Gong, Yehao Li, Tao Mei, Yingwei Pan, Yi Peng, Zhaofan Qiu, Ting Yao, Kai Yu, Yiheng Zhang, et al. Hidream-o1-image: A natively unified image generative foundation model with pixel-level unified transformer.arXiv preprint arXiv:2605.1...

  10. [18]

    Rwku: Benchmarking real-world knowledge unlearning for large language models.Advances in Neural Information Processing Systems, 37:98213–98263, 2024

    Pengfei Cao, Chenhao Wang, Zhitao He, Hongbang Yuan, Jiachun Li, Yubo Chen, Kang Liu, Jun Zhao, et al. Rwku: Benchmarking real-world knowledge unlearning for large language models.Advances in Neural Information Processing Systems, 37:98213–98263, 2024

  11. [19]

    Hunyuanimage 3.0 technical report.arXiv preprint arXiv:2509.23951, 2025

    Siyu Cao, Hangting Chen, Peng Chen, Yiji Cheng, Yutao Cui, Xinchi Deng, Ying Dong, Kipper Gong, Tianpeng Gu, Xiusen Gu, et al. Hunyuanimage 3.0 technical report.arXiv preprint arXiv:2509.23951, 2025

  12. [20]

    Extracting training data from diffusion models

    Nicolas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramer, Borja Balle, Daphne Ippolito, and Eric Wallace. Extracting training data from diffusion models. In32nd USENIX security symposium (USENIX Security 23), pp. 5253–5270, 2023

  13. [21]

    Oneig-bench: Omni-dimensional nuanced evaluation for image generation, 2025

    Jingjing Chang, Yixiao Fang, Peng Xing, Shuhan Wu, Wei Cheng, Rui Wang, Xianfang Zeng, Gang Yu, and Hai-Bao Chen. Oneig-bench: Omni-dimensional nuanced evaluation for image generation, 2025. URL https://arxiv.org/abs/2506.07977

  14. [22]

    Lens: Rethinking training efficiency for foundational text-to-image models

    Dong Chen, Fangyun Wei, Ziyu Wan, Dongdong Chen, Jiawei Zhang, Jinjing Zhao, Sirui Zhang, Yang Yue, Zhiyang Liang, Baining Guo, et al. Lens: Rethinking training efficiency for foundational text-to-image models. arXiv preprint arXiv:2605.21573, 2026

  15. [23]

    Blip3o-next: Next frontier of native image generation, 2025

    Jiuhai Chen, Le Xue, Zhiyang Xu, Xichen Pan, Shusheng Yang, Can Qin, An Yan, Honglu Zhou, Zeyuan Chen, Lifu Huang, Tianyi Zhou, Junnan Li, Silvio Savarese, Caiming Xiong, and Ran Xu. Blip3o-next: Next frontier of native image generation, 2025. URLhttps://arxiv.org/abs/2510.15857

  16. [24]

    Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis, 2023

    Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis, 2023

  17. [25]

    Frugalgpt: How to use large language models while reducing cost and improving performance, 2023

    Lingjiao Chen, Matei Zaharia, and James Zou. Frugalgpt: How to use large language models while reducing cost and improving performance, 2023. URLhttps://arxiv.org/abs/2305.05176

  18. [26]

    Reproducible vision-language models meet concepts out of pre-training

    Ziliang Chen, Xin Huang, Xiaoxuan Fan, Keze Wang, Yuyu Zhou, Quanlong Guan, and Liang Lin. Reproducible vision-language models meet concepts out of pre-training. InProceedings of the Computer Vision and Pattern Recognition Conference, pp. 14701–14711, 2025

  19. [28]

    Goodhart’s law: its origins, meaning and implications for monetary policy

    K Alec Chrystal and Paul D Mizen. Goodhart’s law: its origins, meaning and implications for monetary policy. Central banking, monetary theory and practice: Essays in honour of Charles Goodhart, 1:221–243, 2003

  20. [29]

    Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, Rodolphe Jenatton, Lucas Beyer, Michael Tschannen, Anurag Arnab, Xiao Wang, Carlos Riquelme, Matthias Mindere...

  21. [30]

    Imagenet: A large-scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009

  22. [31]

    Hybrid llm: Cost-efficient and quality-aware query routing

    Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Rühle, Laks Lakshmanan, and Ahmed H Awadallah. Hybrid llm: Cost-efficient and quality-aware query routing. InInternationalConference on Learning Representations, volume 2024, pp. 41348–41366, 2024

  23. [32]

    Cogview: Mastering text-to-image generation via transformers

    Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, and Jie Tang. Cogview: Mastering text-to-image generation via transformers. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (eds.),A...

  24. [33]

    Documenting large webtext corpora: A case study on the colossal clean crawled corpus

    Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 conference on empirical methods in na...

  25. [34]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at...

  26. [35]

    Textcrafter: Accurately rendering multiple texts in complex visual scenes.arXiv e-prints, pp

    Nikai Du, Zhennan Chen, Zhizhou Chen, Shan Gao, Xi Chen, Zhengkai Jiang, Jian Yang, and Ying Tai. Textcrafter: Accurately rendering multiple texts in complex visual scenes.arXiv e-prints, pp. arXiv–2503, 2025

  27. [36]

    Scaling rectified flow transformers for high-resolution image synthesis

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, and Robin Rombach. Scaling rectified flow transformers for high-resolution image synthes...

  28. [37]

    Datacomp: in search of the next generation of multimodal datasets

    Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, Eyal Orgad, Rahim Entezari, Giannis Daras, Sarah Pratt, Vivek Ramanujan, Yonatan Bitton, Kalyani Marathe, Stephen Mussman...

  29. [38]

    Lumina-t2x: Scalable flow-based large diffusion transformer for flexible resolution generation

    Peng Gao, Le Zhuo, Dongyang Liu, Ruoyi Du, Xu Luo, Longtian Qiu, Yuhang Zhang, Rongjie Huang, Shijie Geng, Renrui Zhang, Junlin Xie, Wenqi Shao, Zhengkai Jiang, Tianshuo Yang, Weicai Ye, Tong He, Jingwen He, Junjun He, Yu Qiao, and Hongsheng Li. Lumina-t2x: Scalable flow-based...

  30. [39]

    Seedream 3.0 technical report, 2025

    Yu Gao, Lixue Gong, Qiushan Guo, Xiaoxia Hou, Zhichao Lai, Fanshi Li, Liang Li, Xiaochen Lian, Chao Liao, Liyang Liu, Wei Liu, Yichun Shi, Shiqi Sun, Yu Tian, Zhi Tian, Peng Wang, Rui Wang, Xuanda Wang, Xun Wang, Ye Wang, Guofeng Wu, Jie Wu, Xin Xia, Xuefeng Xiao, Zhonghua Zha...

  31. [40]

    Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities, 2025

    Google Gemini Team. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities, 2025. URLhttps://arxiv.org/abs/2507.06261

  32. [41]

    X-omni: Reinforcement learning makes discrete autoregressive image generative models great again

    Zigang Geng, Yibing Wang, Yeyao Ma, Chen Li, Yongming Rao, Shuyang Gu, Zhao Zhong, Qinglin Lu, Han Hu, Xiaosong Zhang, et al. X-omni: Reinforcement learning makes discrete autoregressive image generative models great again. arXiv preprint arXiv:2507.22058, 2025

  33. [42]

    Geneval: An object-focused framework for evaluating text-to-image alignment

    Dhruba Ghosh, Hannaneh Hajishirzi, and Ludwig Schmidt. Geneval: An object-focused framework for evaluating text-to-image alignment. Advances in Neural Information Processing Systems, 36:52132–52152, 2023

  34. [43]

    Introducing nano banana pro, 2025

    Google. Introducing nano banana pro, 2025. URL https://blog.google/innovation-and-ai/products/ nano-banana-pro/

  35. [44]

    Nano banana 2: Combining pro capabilities with lightning-fast speed, 2026

    Google. Nano banana 2: Combining pro capabilities with lightning-fast speed, 2026. URLhttps://blog.google/ innovation-and-ai/technology/ai/nano-banana-2/

  36. [45]

    Imagen 4.0.https://deepmind.google/models/imagen/, 2025

    Google DeepMind. Imagen 4.0.https://deepmind.google/models/imagen/, 2025

  37. [46]

    Gemini 3.1 pro - model card

    Google DeepMind. Gemini 3.1 pro - model card. https://deepmind.google/models/model-cards/ gemini-3-1-pro/, February 2026. Accessed: 2026-05-05

  38. [47]

    Accurate, large minibatch sgd: Training imagenet in 1 hour, 2018

    Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. Accurate, large minibatch sgd: Training imagenet in 1 hour, 2018. URL https://arxiv.org/abs/1706.02677. 48

  39. [48]

    Accelerate: Training and inference at scale made simple, efficient and adaptable

    Sylvain Gugger, Lysandre Debut, Thomas Wolf, Philipp Schmid, Zachary Mueller, Sourab Mangrulkar, Marc Sun, and Benjamin Bossan. Accelerate: Training and inference at scale made simple, efficient and adaptable. https://github.com/huggingface/accelerate, 2022

  40. [49]

    R-bench: Graduate-level multi-disciplinary benchmarks for llm & mllm complex reasoning evaluation.arXiv preprint arXiv:2505.02018, 2025

    Meng-Hao Guo, Jiajun Xu, Yi Zhang, Jiaxi Song, Haoyang Peng, Yi-Xuan Deng, Xinzhi Dong, Kiyohiro Nakayama, Zhengyang Geng, Chen Wang, et al. R-bench: Graduate-level multi-disciplinary benchmarks for llm & mllm complex reasoning evaluation.arXiv preprint arXiv:2505.02018, 2025

  41. [50]

    Shampoo: Preconditioned stochastic tensor optimization, 2018

    Vineet Gupta, Tomer Koren, and Yoram Singer. Shampoo: Preconditioned stochastic tensor optimization, 2018. URLhttps://arxiv.org/abs/1802.09568

  42. [51]

    Optimizing prompts for text-to-image generation.Advances in Neural Information Processing Systems, 36:66923–66939, 2023

    Yaru Hao, Zewen Chi, Li Dong, and Furu Wei. Optimizing prompts for text-to-image generation.Advances in Neural Information Processing Systems, 36:66923–66939, 2023

  43. [52]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016

  44. [53]

    Gems: Agent-native multimodal generation with memory and skills.arXiv preprint arXiv:2603.28088, 2026

    Zefeng He, Siyuan Huang, Xiaoye Qu, Yafu Li, Tong Zhu, Yu Cheng, and Yang Yang. Gems: Agent-native multimodal generation with memory and skills.arXiv preprint arXiv:2603.28088, 2026

  45. [54]

    Clipscore: A reference-free evaluation metric for image captioning

    Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation metric for image captioning. InProceedings of the 2021 conference on empirical methods in natural language processing, pp. 7514–7528, 2021

  46. [55]

    Classifier-free diffusion guidance

    Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. InNeurIPS2021WorkshoponDeepGenerative Models and Downstream Applications, 2021. URLhttps://openreview.net/forum?id=qw8AKxfYbI

  47. [56]

    Denoising diffusion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. InProceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20, Red Hook, NY, USA, 2020. Curran Associates Inc. ISBN 9781713829546

  48. [57]

    Ella: Equip diffusion models with llm for enhanced semantic alignment.arXiv preprint arXiv:2403.05135, 2024

    Xiwei Hu, Rui Wang, Yixiao Fang, Bin Fu, Pei Cheng, and Gang Yu. Ella: Equip diffusion models with llm for enhanced semantic alignment.arXiv preprint arXiv:2403.05135, 2024

  49. [58]

    Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering

    Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Ostendorf, Ranjay Krishna, and Noah A Smith. Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering. InProceedings of the IEEE/CVF International Conference on Computer Vision, pp. 204...

  50. [59]

    T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.Advancesin Neural Information Processing Systems, 36: 78723–78747, 2023

    Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu. T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.Advancesin Neural Information Processing Systems, 36: 78723–78747, 2023

  51. [60]

    Ape: Agentic prompt enhancer for image generation and editing.arXiv preprint arXiv:2606.00204, 2026

    Zijian Huang, Jay Zhangjie Wu, Zian Wang, Tianshi Cao, Jiasi Chen, Sanja Fidler, Huan Ling, and Xuanchi Ren. Ape: Agentic prompt enhancer for image generation and editing.arXiv preprint arXiv:2606.00204, 2026

  52. [61]

    Scaling up visual and vision-language representation learning with noisy text supervision

    Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International conference on machine learning, pp. 4904–4916. P...

  53. [62]

    Geditbench v2: A human-aligned benchmark for general image editing.arXiv preprint arXiv:2603.28547, 2026

    Zhangqi Jiang, Zheng Sun, Xianfang Zeng, Yufeng Yang, Xuanyang Zhang, Yongliang Wu, Wei Cheng, Gang Yu, Xu Yang, and Bihan Wen. Geditbench v2: A human-aligned benchmark for general image editing.arXiv preprint arXiv:2603.28547, 2026

  54. [63]

    Muon: An optimizer for hidden layers in neural networks, 2024

    Keller Jordan, Yuchen Jin, Vlado Boza, Jiacheng You, Franz Cesista, Laker Newhouse, and Jeremy Bernstein. Muon: An optimizer for hidden layers in neural networks, 2024. URLhttps://kellerjordan.github.io/posts/ muon/

  55. [64]

    Scaling laws for neural language models

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020

  56. [65]

    Elucidating the design space of diffusion-based generative models

    Tero Karras, Miika Aittala, Samuli Laine, and Timo Aila. Elucidating the design space of diffusion-based generative models. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22, Red Hook, NY, USA, 2022. Curran Associates Inc. I...

  57. [66]

    Analyzing and improving the training dynamics of diffusion models, 2024

    Tero Karras, Miika Aittala, Jaakko Lehtinen, Janne Hellsten, Timo Aila, and Samuli Laine. Analyzing and improving the training dynamics of diffusion models, 2024. URLhttps://arxiv.org/abs/2312.02696

  58. [67]

    Auto-encoding variational bayes.arXiv preprint arXiv:1312.6114, 2013

    Diederik P Kingma and Max Welling. Auto-encoding variational bayes.arXiv preprint arXiv:1312.6114, 2013

  59. [68]

    Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advancesin neural information processing systems, 36:36652–36663, 2023

    Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advancesin neural information processing systems, 36:36652–36663, 2023

  60. [69]

    Imagenet classification with deep convolutional neural networks

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012

  61. [70]

    Kling-image-2.1.https://kling.ai/explore/kling_2.1_api, 2025

    Kuaishou Technology. Kling-image-2.1.https://kling.ai/explore/kling_2.1_api, 2025

  62. [71]

    Improved precision and recall metric for assessing generative models.Advances in neural information processing systems, 32, 2019

    Tuomas Kynkäänniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, and Timo Aila. Improved precision and recall metric for assessing generative models.Advances in neural information processing systems, 32, 2019

  63. [72]

    Krea 2.https://www.krea.ai/blog/krea-2-technical-report, 2026

    Sangwu Lee, Erwann Millon, Le Zhuo, Matthew Newton, Andrei Filatov, Naga Sai Abhinay Devarinti, Dazhi Zhong, Avram Djordjevic, Gabriel Menezes, Will Beddow, Titus Ebbecke, Mihai Petrescu, Owen Fahey, Gian Saß, Felix Gil, and Victor Perez. Krea 2.https://www.krea.ai/blog/krea-2...

  64. [73]

    Qwen-image-bench: From generation to creation in text-to-image evaluation,

    Niantong Li, Guangzheng Hu, Weixu Qiao, Ying Ba, Qichen Hong, Shijun Shen, Jinlin Wang, Fan Zhou, Jianye Kang, Xin Shang, Ziyi He, Wei Wang, Dalin Li, Jiahao Li, Jie Zhang, Kaiyuan Gao, Kun Yan, Lihan Jiang, Ningyuan Tang, Shengming Yin, Tianhe Wu, Xiao Xu, Xiaoyue Chen, Yuxia...

  65. [74]

    Omnicorpus: A unified multimodal corpus of 10 billion-level images interleaved with text

    Qingyun Li, Zhe Chen, Weiyun Wang, Wenhai Wang, Shenglong Ye, Zhenjiang Jin, Guanzhou Chen, Yinan He, Zhangwei Gao, Erfei Cui, Jiashuo Yu, Hao Tian, Jiasheng Zhou, Chao Xu, Bin Wang, Xingjian Wei, Wei Li, Wenjian Zhang, Bo Zhang, Pinlong Cai, Licheng Wen, Xiangchao Yan, Pei Ch...

  66. [75]

    Reflect-dit: Inference-time scaling for text-to-image diffusion transformers via in-context reflection

    Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Arsh Koneru, Yusuke Kato, Kazuki Kozuka, and Aditya Grover. Reflect-dit: Inference-time scaling for text-to-image diffusion transformers via in-context reflection. In Proceedings of the IEEE/CVF International Conference on Co...

  67. [76]

    Holistic evaluation of language models.arXiv preprint arXiv:2211.09110, 2022

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. Holistic evaluation of language models.arXiv preprint arXiv:2211.09110, 2022

  68. [77]

    Scaling laws for diffusion transformers.arXiv preprint arXiv:2410.08184, 2024

    Zhengyang Liang, Hao He, Ceyuan Yang, and Bo Dai. Scaling laws for diffusion transformers.arXiv preprint arXiv:2410.08184, 2024

  69. [78]

    Aegis: Exploring the limit of world knowledge capabilities for unified mulitmodal models.arXiv preprint arXiv:2601.00561, 2026

    Jintao Lin, Bowen Dong, Weikang Shi, Chenyang Lei, Suiyun Zhang, Rui Liu, and Xihui Liu. Aegis: Exploring the limit of world knowledge capabilities for unified mulitmodal models.arXiv preprint arXiv:2601.00561, 2026

  70. [79]

    Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling, 2023. URLhttps://arxiv.org/abs/2210.02747

  71. [80]

    Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

  72. [81]

    Ernie-image technical report.arXiv preprint arXiv:2605.25347, 2026

    Jiaxiang Liu, Zhida Feng, Pengyu Zou, Zhenyu Qian, Tianrui Zhu, Jun Xia, Yuehu Dong, Yanzheng Lin, Honglin Xiong, Anqi Chen, et al. Ernie-image technical report.arXiv preprint arXiv:2605.25347, 2026

  73. [82]

    Flow-grpo: Training flow matching models via online rl, 2025

    Jie Liu, Gongye Liu, Jiajun Liang, Yangguang Li, Jiaheng Liu, Xintao Wang, Pengfei Wan, Di Zhang, and Wanli Ouyang. Flow-grpo: Training flow matching models via online rl, 2025. URLhttps://arxiv.org/abs/2505. 05470

  74. [83]

    Step1x-edit: Apracticalframeworkforgeneralimageediting

    Shiyu Liu, Yucheng Han, Peng Xing, Fukun Yin, Rui Wang, Wei Cheng, Jiaqi Liao, Yingming Wang, Honghao Fu, ChunruiHan, etal. Step1x-edit: Apracticalframeworkforgeneralimageediting. arXivpreprintarXiv:2504.17761, 2025. 50

  75. [84]

    Flow straight and fast: Learning to generate and transfer data with rectified flow, 2022

    Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow, 2022. URLhttps://arxiv.org/abs/2209.03003

  76. [85]

    Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps, 2022

    Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps, 2022. URLhttps://arxiv.org/abs/2206.00927

  77. [86]

    Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models.Machine Intelligence Research, 22(4):730–751, June 2025

    Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models.Machine Intelligence Research, 22(4):730–751, June 2025. ISSN 2731-5398. doi: 10.1007/s11633-025-1562-4. URLhttp://dx.doi.org...

  78. [87]

    Ovis: Structural embedding alignment for multimodal large language model.arXiv preprint arXiv:2405.20797, 2024

    Shiyin Lu, Yang Li, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, and Han-Jia Ye. Ovis: Structural embedding alignment for multimodal large language model.arXiv preprint arXiv:2405.20797, 2024

  79. [88]

    Murphy.Probabilistic Machine Learning: Advanced Topics

    Kevin P. Murphy.Probabilistic Machine Learning: Advanced Topics. MIT Press, 2023. URLhttp://probml. github.io/book2

  80. [89]

    Some matrix-inequalities and metrization of matric-space.Tomsk.Univ

    Johann von Neumann. Some matrix-inequalities and metrization of matric-space.Tomsk.Univ. Rev., 1:286–300,

  81. [90]

    Improveddenoisingdiffusionprobabilisticmodels

    AlexanderQuinnNicholandPrafullaDhariwal. Improveddenoisingdiffusionprobabilisticmodels. In International conference on machine learning, pp. 8162–8171. PMLR, 2021

  82. [91]

    Wise: A world knowledge-informed semantic evaluation for text-to-image generation

    Yuwei Niu, Munan Ning, Mengren Zheng, Weiyang Jin, Bin Lin, Peng Jin, Jiaqi Liao, Kunpeng Ning, Chaoran Feng, Bin Zhu, and Li Yuan. Wise: A world knowledge-informed semantic evaluation for text-to-image generation. arXiv preprint arXiv:2503.07265, 2025

  83. [92]

    Routellm: Learning to route llms with preference data.arXiv preprint arXiv:2406.18665, 2024

    Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E Gonzalez, M Waleed Kadous, and Ion Stoica. Routellm: Learning to route llms with preference data.arXiv preprint arXiv:2406.18665, 2024

  84. [93]

    Gpt image 1.https://developers.openai.com/api/docs/models/gpt-image-1, 2025

    OpenAI. Gpt image 1.https://developers.openai.com/api/docs/models/gpt-image-1, 2025

  85. [94]

    Introducing chatgpt images 2.0, 2026

    OpenAI. Introducing chatgpt images 2.0, 2026. URL https://openai.com/index/ introducing-chatgpt-images-2-0/

  86. [95]

    The neglected tails in vision-language models, 2024

    Shubham Parashar, Zhiqiu Lin, Tian Liu, Xiangjue Dong, Yanan Li, Deva Ramanan, James Caverlee, and Shu Kong. The neglected tails in vision-language models, 2024. URLhttps://arxiv.org/abs/2401.12425

  87. [96]

    On aliased resizing and surprising subtleties in gan evaluation

    Gaurav Parmar, Richard Zhang, and Jun-Yan Zhu. On aliased resizing and surprising subtleties in gan evaluation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11410–11420, 2022

  88. [97]

    Scalable diffusion models with transformers, 2023

    William Peebles and Saining Xie. Scalable diffusion models with transformers, 2023. URLhttps://arxiv.org/ abs/2212.09748

  89. [98]

    Sdxl: Improving latent diffusion models for high-resolution image synthesis

    Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis. InInternational Conference on Learning Representations, volume 2024, pp. 1862–1874, 2024

  90. [99]

    Qwen-image-2512: Finer details, greater realism.https://qwen.ai/blog?id= qwen-image-2512, 2025

    Qwen Team, Alibaba Group. Qwen-image-2512: Finer details, greater realism.https://qwen.ai/blog?id= qwen-image-2512, 2025. Apache-2.0 License

  91. [100]

    Qwen-image-edit-2509: Multi-image support, improved consistency

    Qwen Team, Alibaba Group. Qwen-image-edit-2509: Multi-image support, improved consistency. https: //qwen.ai/blog?id=7a90090115ee193ce6a7f619522771dd9696dd93, 2025. Apache-2.0 License

  92. [101]

    Qwen-image-2511.https://qwen.ai/blog?id=qwen-image-2512, 2025

    Qwen Team, Alibaba Group. Qwen-image-2511.https://qwen.ai/blog?id=qwen-image-2512, 2025. Apache-2.0 License

  93. [102]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp...

  94. [103]

    Zero: Memory optimizations toward training trillion parameter models, 2020

    Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. Zero: Memory optimizations toward training trillion parameter models, 2020. URLhttps://arxiv.org/abs/1910.02054

  95. [104]

    Hierarchical text-conditional image generation with clip latents.arXiv preprint arXiv:2204.06125, 1(2):3, 2022

    Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents.arXiv preprint arXiv:2204.06125, 1(2):3, 2022. 51

  96. [105]

    Do imagenet classifiers generalize to imagenet? In International conference on machine learning, pp

    Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do imagenet classifiers generalize to imagenet? In International conference on machine learning, pp. 5389–5400. PMLR, 2019

  97. [106]

    A structured review of the validity of bleu.Computational Linguistics, 44(3):393–401, 2018

    Ehud Reiter. A structured review of the validity of bleu.Computational Linguistics, 44(3):393–401, 2018

  98. [107]

    High-resolution image synthesis with latent diffusion models, 2022

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models, 2022. URLhttps://arxiv.org/abs/2112.10752

  99. [108]

    U-net: Convolutional networks for biomedical image segmentation

    Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 234–241. Springer, 2015

  100. [109]

    Seyedmorteza Sadat, Otmar Hilliges, and Romann M. Weber. Eliminating oversaturation and artifacts of high guidance scales in diffusion models, 2025. URLhttps://arxiv.org/abs/2410.02416

  101. [110]

    Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information ...

  102. [111]

    Laion-400m: Open dataset of clip-filtered 400 million image-text pairs

    Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021

  103. [112]

    Laion-5b: An open large-scale dataset for training next generation image-text models.Advancesin neural information processing systems, 35:25278–25294, 2022

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models.Advancesin neural informati...

  104. [113]

    Seedream 4.0: Toward next-generation multimodal image generation, 2025

    Team Seedream, :, Yunpeng Chen, Yu Gao, Lixue Gong, Meng Guo, Qiushan Guo, Zhiyao Guo, Xiaoxia Hou, Weilin Huang, Yixuan Huang, Xiaowen Jian, Huafeng Kuang, Zhichao Lai, Fanshi Li, Liang Li, Xiaochen Lian, Chao Liao, Liyang Liu, Wei Liu, Yanzuo Lu, Zhengxiong Luo, Tongtong Ou,...

  105. [114]

    From pixels to prose: A large dataset of dense image captions, 2024

    Vasu Singla, Kaiyu Yue, Sukriti Paul, Reza Shirkavand, Mayuka Jayawardhana, Alireza Ganjdanesh, Heng Huang, Abhinav Bhatele, Gowthami Somepalli, and Tom Goldstein. From pixels to prose: A large dataset of dense image captions, 2024. URLhttps://arxiv.org/abs/2406.10328

  106. [115]

    Weiss, Niru Maheswaranathan, and Surya Ganguli

    Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics, 2015. URLhttps://arxiv.org/abs/1503.03585

  107. [116]

    Diffusion art or digital forgery? investigating data replication in diffusion models

    Gowthami Somepalli, Vasu Singla, Micah Goldblum, Jonas Geiping, and Tom Goldstein. Diffusion art or digital forgery? investigating data replication in diffusion models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 6048–6058, 2023

  108. [117]

    Denoising diffusion implicit models, 2022

    Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models, 2022. URLhttps: //arxiv.org/abs/2010.02502

  109. [118]

    Joyai-image: Awaking spatial intelligence in unified multimodal understanding and generation

    Lin Song, Wenbo Li, Guoqing Ma, Wei Tang, Bo Wang, Yuan Zhang, Yijun Yang, Yicheng Xiao, Jianhui Liu, Yanbing Zhang, et al. Joyai-image: Awaking spatial intelligence in unified multimodal understanding and generation. arXiv preprint arXiv:2605.04128, 2026

  110. [119]

    Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole

    Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations, 2021. URLhttps://arxiv.org/abs/ 2011.13456

  111. [120]

    Introducing stable diffusion 3.5

    Stability AI. Introducing stable diffusion 3.5. https://stability.ai/news/ introducing-stable-diffusion-3-5, Oct 2024

  112. [121]

    Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 2024

    Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 2024. 52

  113. [122]

    Longcat-image technical report.arXiv preprint arXiv:2512.07584, 2025

    Meituan LongCat Team, Hanghang Ma, Haoxian Tan, Jiale Huang, Junqiang Wu, Jun-Yan He, Lishuai Gao, Songlin Xiao, Xiaoming Wei, Xiaoqi Ma, et al. Longcat-image technical report.arXiv preprint arXiv:2512.07584, 2025

  114. [123]

    Firered-image-edit-1.0 technical report.arXiv preprint arXiv:2602.13344, 2026

    Super Intelligence Team, Changhao Qiao, Chao Hui, Chen Li, Cunzheng Wang, Dejia Song, Jiale Zhang, Jing Li, Qiang Xiang, Runqi Wang, et al. Firered-image-edit-1.0 technical report.arXiv preprint arXiv:2602.13344, 2026

  115. [124]

    zero-shot

    Vishaal Udandarao, Ameya Prabhu, Adhiraj Ghosh, Yash Sharma, Philip H. S. Torr, Adel Bibi, Samuel Albanie, and Matthias Bethge. No "zero-shot" without exponential data: Pretraining concept frequency determines multimodal model performance, 2024. URLhttps://arxiv.org/abs/2404.04125

  116. [125]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (eds.), Advances in Neural In...

  117. [126]

    Jco-mvton: Jointly controllable multi-modal diffusion transformer for mask-free virtual try-on.arXiv preprint arXiv:2508.17614, 2025

    Aowen Wang, Wei Li, Hao Luo, Mengxing Ao, Chenyu Zhu, Xinyang Li, and Fan Wang. Jco-mvton: Jointly controllable multi-modal diffusion transformer for mask-free virtual try-on.arXiv preprint arXiv:2508.17614, 2025

  118. [127]

    Ovis-u1 technical report, 2025

    Guo-Hua Wang, Shanshan Zhao, Xinjie Zhang, Liangfu Cao, Pengxin Zhan, Lunhao Duan, Shiyin Lu, Minghao Fu, Xiaohao Chen, Jianshan Zhao, Yang Li, and Qing-Guo Chen. Ovis-u1 technical report, 2025. URL https://arxiv.org/abs/2506.23044

  119. [128]

    Pytorch image models.https://github.com/rwightman/pytorch-image-models, 2019

    Ross Wightman. Pytorch image models.https://github.com/rwightman/pytorch-image-models, 2019

  120. [129]

    Resnet strikes back: An improved training procedure in timm, 2021

    Ross Wightman, Hugo Touvron, and Hervé Jégou. Resnet strikes back: An improved training procedure in timm, 2021. URLhttps://arxiv.org/abs/2110.00476

  121. [130]

    Qwen-image technical report, 2025

    Chenfei Wu, Jiahao Li, Jingren Zhou, Junyang Lin, Kaiyuan Gao, Kun Yan, Sheng ming Yin, Shuai Bai, Xiao Xu, Yilei Chen, Yuxiang Chen, Zecheng Tang, Zekai Zhang, Zhengyi Wang, An Yang, Bowen Yu, Chen Cheng, Dayiheng Liu, Deqing Li, Hang Zhang, Hao Meng, Hu Wei, Jingyuan Ni, Kai...

  122. [131]

    Omnigen2: Exploration to advanced multimodal generation, 2025

    Chenyuan Wu, Pengfei Zheng, Ruiran Yan, Shitao Xiao, Xin Luo, Yueze Wang, Wanli Li, Xiyan Jiang, Yexin Liu, Junjie Zhou, Ze Liu, Ziyi Xia, Chaofan Li, Haoge Deng, Jiahao Wang, Kun Luo, Bo Zhang, Defu Lian, Xinlong Wang, Zhongyuan Wang, Tiejun Huang, and Zheng Liu. Omnigen2: Ex...

  123. [132]

    Omnigen: Unified image generation

    Shitao Xiao, Yueze Wang, Junjie Zhou, Huaying Yuan, Xingrun Xing, Ruiran Yan, Chaofan Li, Shuting Wang, Tiejun Huang, and Zheng Liu. Omnigen: Unified image generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13294–13304, 2025

  124. [133]

    Sana: Efficient high-resolution image synthesis with linear diffusion transformer,

    Enze Xie, Junsong Chen, Junyu Chen, Han Cai, Haotian Tang, Yujun Lin, Zhekai Zhang, Muyang Li, Ligeng Zhu, Yao Lu, and Song Han. Sana: Efficient high-resolution image synthesis with linear diffusion transformer,

  125. [134]

    Imagereward: Learning and evaluating human preferences for text-to-image generation

    JiazhengXu, XiaoLiu, YuchenWu, YuxuanTong, QinkaiLi, MingDing, JieTang, andYuxiaoDong. Imagereward: Learning and evaluating human preferences for text-to-image generation. Advances in Neural Information Processing Systems, 36:15903–15935, 2023

  126. [135]

    Qwen3 technical report, 2025

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jia...

  127. [136]

    Imgedit: A unified image editing dataset and benchmark.Advances in Neural Information Processing Systems, 38, 2026

    Yang Ye, Xianyi He, Zongjian Li, Shenghai Yuan, Zhiyuan Yan, Bohan Hou, Li Yuan, et al. Imgedit: A unified image editing dataset and benchmark.Advances in Neural Information Processing Systems, 38, 2026. 53

  128. [137]

    Scaling autoregressive models for content-rich text-to-image generation

    Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al. Scaling autoregressive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789, 2(3):5, 2022

  129. [138]

    Z-image: An efficient image generation foundation model with single-stream diffusion transformer, 2025

    Z-Image Team, Huanqia Cai, Sihan Cao, Ruoyi Du, Peng Gao, Steven Hoi, Zhaohui Hou, Shijie Huang, Dengyang Jiang, Xin Jin, Liangchen Li, Zhen Li, Zhong-Yu Li, David Liu, Dongyang Liu, Junhan Shi, Qilong Wu, Feng Yu, Chi Zhang, Shifeng Zhang, and Shilin Zhou. Z-image: An efficie...

  130. [139]

    Glm-image: Auto-regressive for dense-knowledge and high-fidelity image generation.https://z.ai/blog/ glm-image, 2026

    Z.AI. Glm-image: Auto-regressive for dense-knowledge and high-fidelity image generation.https://z.ai/blog/ glm-image, 2026. Apache-2.0 License

  131. [140]

    Worldgenbench: A world-knowledge-integrated benchmark for reasoning-driven text-to-image generation

    Daoan Zhang, Che Jiang, Ruoshi Xu, Biaoxiang Chen, Zijian Jin, Yutian Lu, Jianguo Zhang, Liang Yong, Jiebo Luo, and Shengda Luo. Worldgenbench: A world-knowledge-integrated benchmark for reasoning-driven text-to-image generation. arXiv preprint arXiv:2505.01490, 2025

  132. [141]

    Adding conditional control to text-to-image diffusion models

    Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 3836–3847, 2023

  133. [142]

    Qwen-image-agent: Bridging the context gap in real-world image generation, 2026

    Zekai Zhang, Jiahao Li, Jie Zhang, Kaiyuan Gao, Kun Yan, Lihan Jiang, Ningyuan Tang, Shengming Yin, Tianhe Wu, Xiaoyue Chen, Xiao Xu, Yan Shu, Yanran Zhang, Yixian Xu, Yuxiang Chen, Zhendong Wang, Zihao Liu, Zikai Zhou, Huishuai Zhang, Dongyan Zhao, and Chenfei Wu. Qwen-image-...

  134. [143]

    Qwen-image-2.0 technical report.arXiv preprint arXiv:2605.10730, 2026

    Bing Zhao, Chenfei Wu, Deqing Li, Hao Meng, Jiahao Li, Jie Zhang, Jingren Zhou, Junyang Lin, Kaiyuan Gao, Kuan Cao, et al. Qwen-image-2.0 technical report.arXiv preprint arXiv:2605.10730, 2026

  135. [144]

    Pytorch fsdp: Experiences on scaling fully sharded data parallel, 2023

    Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews, and Shen Li. Pytorch fsdp: Experiences on sc...

  136. [145]

    Chatbot arena: An open platform for evaluating llms based on human preference

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhi Zheng Lin, Zi Ning Li, Dacheng Li, Eric Xing, et al. Chatbot arena: An open platform for evaluating llms based on human preference. In arXiv preprint arXiv:2403.04132, 2024

  137. [146]

    Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025

    Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, Zhangwei Gao, Erfei Cui, Xuehui Wang, Yue Cao, Yangzhou Liu, Xingguang Wei, Hongjie Zhang, Haomin Wang, Weiye Xu, Hao Li, Jiahao Wang, Nianchen Deng, Songze Li,...

  138. [147]

    Lumina-next: making lumina-t2x stronger and faster with next-dit

    Le Zhuo, Ruoyi Du, Han Xiao, Yangguang Li, Dongyang Liu, Rongjie Huang, Wenze Liu, Xiangyang Zhu, Fu-Yun Wang, Zhanyu Ma, Xu Luo, Zehan Wang, Kaipeng Zhang, Lirui Zhao, Si Liu, Xiangyu Yue, Wanli Ouyang, Yu Qiao, Hongsheng Li, and Peng Gao. Lumina-next: making lumina-t2x stron...

  139. [153]

    and Z-Image [138] integrate deep visual understanding with scalable, efficient generation, paving the way for versatile and intuitive multimodal AI systems. B.2 Macro-Architecture Design Empirical observations of state-of-the-art proprietary models, such as ChatGPT Images seri...

  140. [154]

    film-like

    Instruction Encoder Qwen3-VL-8B-Instruct [3] Prompt Tuning (Optional) Trainable Tokens 32 Causal Mask True Hidden Size 4096 Layers 3 Attention Heads 32 KV Heads 8 Trainable Parameters DiT Model 10,292.56M Total (w/ Prompt Tuning) 11,022.55M Table 10Summary of architectural des...

  141. [1937]

    Reprinted inJohn von Neumann: Collected Works(A. H. Taub, ed.), Vol. IV, Pergamon Press, Oxford, 1962, pp. 205-218

  142. [2023]

    URLhttps://arxiv.org/abs/2302.05442

  143. [2024]

    URLhttps://arxiv.org/abs/2410.10629

  144. [2025]

    URLhttps://bfl.ai/research/representation-comparison

  145. [2026]

    URLhttps://arxiv.org/abs/2605.28091

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.