Pith. sign in

REVIEW 4 major objections 5 minor 153 references

Boogu-Image-0.1 argues that upgrading the understanding side of a text-to-image system—encoder, prompt rewriting, captions, routing—can make a model trained on 208.62M images at roughly $400K competitive with far costlier open- and closed-s

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 06:07 UTC pith:2JDMA2ZH

load-bearing objection Genuinely useful engineering report with honest ablations; flagship SOTA claim rests on an in-house benchmark needing external validation. the 4 major comments →

arxiv 2607.13125 v2 pith:2JDMA2ZH submitted 2026-07-14 cs.CV cs.AI

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget

classification cs.CV cs.AI
keywords text-to-image generationimage editingmultimodal understandingagentic prompt rewritingtraining efficiencydata curationElo evaluationbilingual text rendering
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper's central claim is that image-generation quality is gated less by generator size or data volume than by how well the system understands instructions. It builds a unified understanding-and-generation family around a stronger instruction encoder, an agentic prompt rewriter, captions designed around user demand, and a model router, and reports that this lets a from-scratch 10B-parameter model trained on only 208.62M unique images match or beat other open models and approach closed frontier systems. On its own in-house blind-vote benchmark and on a post-freeze open benchmark, the models lead the open-source tier, with the thinking variant adding the largest gains in text rendering and creativity. The authors also report practical data findings: a curated 47M-image syllabus beats 187M raw open-source images, roughly 300 exposures reliably teach a Chinese glyph, and capping the timestep-shift at the 1K token-count regime fixes slow convergence at 2K. The authors explicitly caution that their primary arena is in-house and that the editing benchmark inverts under human evaluation, so the headline rankings rest on instruments that await independent validation.

Core claim

On its own terms, the paper establishes that understanding can be engineered as a first-class, decoupled component of a generation system rather than left implicit in a large generator. It decomposes understanding into an instruction encoder that acts as a lossy sensor, an agentic prompt rewriter that translates intent without embellishing clear prompts, captions that name what users will ask for (including visual flaws), and a difficulty-aware router that dispatches easy requests to fast models. The empirical core is a series of controlled comparisons: encoder scale alone moves GenEval from 0.60 to 0.65 with no saturation; the same 47M-image syllabus outperforms 187M of open-source data; th

What carries the argument

The load-bearing mechanism is a decoupled understanding stack placed upstream of generation. A mid-sized frozen vision-language model serves as the instruction encoder, treated as a sensor whose capacity upper-bounds what the diffusion transformer can receive, while a separate reasoning-capable agentic rewriter and a per-aspect captioning pipeline supply the conditioning and supervision. Completing the stack is a router that classifies request difficulty and selects fast versus heavy model variants, turning inference time into a dial on a quality-latency curve. Two supporting techniques carry much of the training-efficiency claim: a syllabus-guided data curriculum with explicit defect captio

Load-bearing premise

The headline top-open-source-model ranking depends on Boogu Arena, an in-house Elo benchmark whose agreement with public human preference was measured on only five models; if that arena does not reflect real user preference, the central competitive claim is unsupported.

What would settle it

Run a public, independent blind pairwise comparison, for example on a held-out bilingual prompt set, between Boogu-Image-0.1-Turbo-Thinking and the leading open baselines it claims to beat, and check whether its Elo still leads; since the weights are released, this is directly executable.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Open-source teams can budget for roughly 200M unique images and about $400K of compute and still reach the top of the open tier, provided understanding-side components are treated as part of the model.
  • Evaluation that reports quality without latency is incomplete: the same Boogu model occupies many points on the quality-time curve, and thinking and routing variants make that curve explicit.
  • Capabilities that are not named in captions cannot be recovered at inference time, so captioning should be planned per anticipated user demand rather than delegated to a single vision-language model.
  • Saturated academic benchmarks such as GenEval and DPG-Bench no longer track human preference, so the paper's results on its arena and on a post-freeze open benchmark carry the evidential weight.
  • Skills embedded in the rewriter, such as counting, infographic layout, reasoning, NSFW handling, and scene text, transfer into measurable generation gains in those same dimensions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the roughly 300-exposure glyph result generalizes, it predicts a cheap and testable recipe for any new writing system or token set: enumerate the vocabulary and floor each unit's exposure rather than scaling raw data.
  • The reported inversion on the editing benchmark, where a closed system scores lower yet wins in the authors' own human evaluation, suggests that model-based editing metrics systematically compress quality gaps; a public human editing arena could settle which ranking is real.
  • The finding that heavy aesthetic reinforcement learning narrows output distributions implies a policy tension: optimizing average appeal can erase off-mainstream demographics and styles, which matters for fairness and access as much as for diversity.
  • Releasing weights with full recipes makes the understanding-first budget claim directly falsifiable and re-runnable, which is the main reason the cost figure can be checked rather than taken on faith.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces Boogu-Image-0.1, an open-source unified text-to-image and image-editing model family (Base, Turbo, Edit, Edit-Turbo), and argues that strengthening the understanding component — via a stronger frozen multimodal encoder (Qwen3-VL-8B), an agentic prompt rewriter, a curated data syllabus, and agentic inference-time scaling — substantially improves generation quality under a minimal compute budget (208.62M unique images, ~$400K theoretical training cost). The evaluation relies on an in-house blind pairwise benchmark (Boogu Arena), on external benchmarks Qwen-Image-Bench and LongText-Bench, and on ImgEdit-Bench for editing, claiming top-tier open-source status and performance approaching closed-source systems.

Significance. If the central claims hold, the paper offers two useful contributions: (i) a systematic demonstration, with ablations, that instruction-encoder scale, curated data syllabus, and agentic prompt rewriting each yield measurable gains under a fixed budget, and (ii) a practical open release of weights, code, and training recipes. The ablations in Tables 7–8 and Figure 18 are genuine and non-circular, comparing actual checkpoints and data configurations. The bottleneck is the headline positioning: the claim of "approaching closed-source systems" and of open-source leadership on human preference rests almost entirely on an in-house benchmark whose external validation is thin, and the paper itself documents an inversion between its automated editing benchmark and its own human evaluation. The external Qwen-Image-Bench results support open-source leadership only among the subset of models evaluated and do not support the closed-source gap claim.

major comments (4)
  1. [§2.2.1, Figure 7] The Boogu Arena Elo ranking is the primary evidence for the headline "approaching closed-source" claim, but its external validation is limited to five models, none of which are Boogu models. With n=5, a Pearson r of 0.986 and Spearman 1.0 are consistent with a linear trend but do not establish that Boogu's absolute Elo position transfers to the LMArena scale; the Boogu points are extrapolations. The paper reports no confidence intervals on Elo, no inter-annotator agreement, and no independent annotation protocol, and the prompts are only promised for future release. The paper itself (Sec. 2.3.1) shows that its in-house preference instrument can invert the ranking of two models relative to human evaluation, so Boogu Arena cannot currently carry the closed-source gap claim.
  2. [§2.3.1, Table 6] The text states that Boogu-Image-0.1-Edit-Thinking "attains the best overall score" on ImgEdit-Bench, but the immediately following Discussion and Limitations says that in the authors' own human evaluations Nano-Banana-Pro outperforms Boogu despite scoring lower (4.37 vs 4.64), and the table is labeled "for reference only." This is not a minor caveat: the abstract and introduction use "matches or surpasses" across standard benchmarks without this qualification, and the internal inversion means the ImgEdit ranking should not be presented as a headline result. The claim needs to be recast or the human evaluation data reported.
  3. [§2.2.2, Tables 1–2; §2.2.3, Table 3] The open-source leadership claim on Qwen-Image-Bench and LongText-Bench is limited by the evaluated set: the text says "due to time constraints, the evaluation does not yet cover all available open-source baselines," and no inference-time budget is reported for any model. Since the paper argues (Sec. 3.2.1) that quality must be reported jointly with latency, the thinking variants, which use agentic rewriting and reflection, should be compared with closed-source systems under a matched or explicitly reported inference budget; otherwise the 'approaching closed-source' portion of the claim is not supported by these tables.
  4. [§3.2.3 (Dynamic time shifting) and Figure 32] The rectified dynamic time-shifting scheme introduces a free parameter ncap (e.g., 4096), and its benefit is illustrated by quantile plots rather than by a controlled training ablation showing that the rectified scheme improves final generation metrics (e.g., FID or benchmark scores) relative to standard dynamic shifting at 2K. The claim that this is a decisive detail for the Boogu pipeline is plausible but currently under-supported; either add an ablation or soften the presentation.
minor comments (5)
  1. [Figure 1] The right panel lists 'Boogu-Image-0.1-Turbo-PE' but the model family in the abstract and Section 2 includes Base, Turbo, Edit, Edit-Turbo; PE is not defined in the caption or text. Please clarify or rename.
  2. [Tables 1–6] Model names are inconsistent: 'Nano-Banana-Pro' vs 'NanoBanana-Pro', and 'Qwen-Image-2.0-2026-03-03' vs 'Qwen-Image-2.0'. Also, some tables say 'Our model series achieves...' while the model is highlighted in red; ensure the caption convention is uniform.
  3. [§2.2.4] The decision to report GenEval and DPG-Bench 'for reference only' is reasonable, but the same standards should apply to ImgEdit-Bench, which is also based on VLM scoring and acknowledged to invert human preference. Consider adding a unified statement about which benchmarks the authors regard as primary and why.
  4. [§4] The limitations paragraph is candid but does not mention the absence of an external human-preference validation for Boogu Arena; this is the most load-bearing limitation for the abstract's claim and should be stated there.
  5. [Abstract] The phrase 'the base model's theoretical training cost is only approximately $400K' is used without defining what 'theoretical' includes (hardware, electricity, ablations, data curation). Please provide a cost breakdown or soften the claim.

Circularity Check

0 steps flagged

No significant circularity: central claims rest on independent ablations and external benchmarks; Boogu Arena is a validity concern, not a circular reduction.

full rationale

The paper's load-bearing derivations are direct empirical ablations over actual checkpoints and data configurations: Table 7 varies the frozen instruction encoder and measures GenEval; Figure 18 scales the rewriter backbone and reports a defined advantage score; Table 8 compares training on 187M open-source data vs. the Boogu Syllabus on Qwen-Image-Bench and LongText-Bench; Figures 27-28 and Table 9 measure glyph-exposure thresholds. In each case the manipulated variable is not defined in terms of the measured outcome, so no self-definitional reduction is present. The supporting benchmarks (Qwen-Image-Bench, LongText-Bench, GenEval, DPG-Bench) are external to the paper. The Boogu Arena leaderboard is built in-house, but its Elo scores are not fitted to Boogu outputs, and the LMArena agreement check (Figure 7), while based on only five non-Boogu models, is a calibration exercise rather than a definitional reduction. The only apparent self-citation ([78], used as an example in the open-challenges discussion) is not load-bearing. Benchmark-validity concerns about Boogu Arena are legitimate evaluation risks, but they are not circularity under the definitions used here.

Axiom & Free-Parameter Ledger

1 free parameters · 4 axioms · 0 invented entities

The central claim rests on a few domain assumptions about evaluation validity and generality of exposure thresholds; no free parameters are fitted to make the derivation work, and no new physical entities are introduced.

free parameters (1)
  • ncap (dynamic time-shift clamp) = 4096
    Chosen by hand to inherit 1K-resolution training behavior for 2K training; not part of the claim-derivation, but an ad hoc training choice reported in Section 3.2.3.
axioms (4)
  • domain assumption The Boogu Arena Elo computed from in-house blind pairwise votes is a faithful proxy for real-world human preference (correlation with LMArena r=0.986 on five anchor models).
    Used to rank Boogu models as top open-source in Section 2.2.1/Figure 6. No inter-annotator agreement or external annotator independence is reported.
  • domain assumption VLM/automated benchmarks (Qwen-Image-Bench, LongText-Bench, ImgEdit-Bench) are meaningful for the claims despite acknowledged saturation and their own ImgEdit inversion.
    Tables 1-6 use these scores as headline evidence; Section 2.3.1 concedes the ImgEdit ranking contradicts human preference.
  • domain assumption A per-exposure threshold (300 per glyph; ~4,500 per identity) measured on a small subset generalizes to the full training distribution.
    Section 3.2.2-3.2.3 derive quantitative claims from ~17 Chinese characters and one volunteer identity.
  • standard math Standard flow-matching/diffusion and DiT formulations are correct as used.
    The training and BOG sections rely on established diffusion/flow-matching math (Sections B and C).

pith-pipeline@v1.3.0-alltime-deepseek · 54065 in / 15435 out tokens · 148266 ms · 2026-08-02T06:07:53.068500+00:00 · methodology

0 comments
read the original abstract

We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual (Chinese-English) text rendering. Closed-source multimodal systems like Nano-Banana-Pro and GPT-Image-2 achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. In this work, we demonstrate that strengthening the understanding capability of the system, through a stronger multimodal encoder, agentic prompt rewriting, and related techniques, together with improvements in data quality, training pipelines, and agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets. Comprehensive evaluations show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks, and achieves results approaching leading closed-source systems. Notably, this is accomplished with only 208.62 million unique images. The base model's theoretical training cost is only approximately \$400K. We share practical discussions that we believe are valuable to the broader research community, and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. Our code is available here: https://github.com/Boogu-Project/Boogu-Image.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

153 extracted references · 60 linked inside Pith

  1. [1]

    Arena ai: The official ai ranking & llm leaderboard.https://arena.ai/, 2026

  2. [2]

    Ideogram 4.https://ideogram.ai/blog/ideogram-4.0/, 2026

    Ideogram AI. Ideogram 4.https://ideogram.ai/blog/ideogram-4.0/, 2026

  3. [3]

    Qwen3-vl technical report, 2025

    Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan ...

  4. [4]

    Qwen2.5-vl technical report, 2025

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report, 2...

  5. [5]

    Beyond the imitation game: Quantifying and extrapolating the capabilities of language models

    BIG bench authors. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URLhttps://openreview.net/ forum?id=uyTL5Bvosj

  6. [6]

    Improving image generation with better captions.Computer Science

    James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al. Improving image generation with better captions.Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3):8, 2023

  7. [7]

    Announcing Black Forest Labs: Black forest labs’ inaugural model suite, FLUX.1, 2024

    Black Forest Labs. Announcing Black Forest Labs: Black forest labs’ inaugural model suite, FLUX.1, 2024. URL https://blackforestlabs.ai/announcing-black-forest-labs/. Accessed: 2026-04-25

  8. [8]

    FLUX.2: Analyzing and enhancing the latent space of FLUX – representation comparison, 11

    Black Forest Labs. FLUX.2: Analyzing and enhancing the latent space of FLUX – representation comparison, 11

  9. [9]

    Flux.2: Production-grade ai image generation and editing model with 4mp photorealistic output and multi-reference control.https://bfl.ai/models/flux-2, 2025

    Black Forest Labs. Flux.2: Production-grade ai image generation and editing model with 4mp photorealistic output and multi-reference control.https://bfl.ai/models/flux-2, 2025

  10. [10]

    Large scale GAN training for high fidelity natural image synthesis

    Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale GAN training for high fidelity natural image synthesis. In International Conference on Learning Representations, 2019. URL https://openreview.net/ forum?id=B1xsqj09Fm

  11. [11]

    Instructpix2pix: Learning to follow image editing instructions

    Tim Brooks, Aleksander Holynski, and Alexei A Efros. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 18392–18402, 2023

  12. [12]

    Video generation models as world simulators.OpenAI Blog, 1(8):1, 2024

    Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Leo Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, et al. Video generation models as world simulators.OpenAI Blog, 1(8):1, 2024

  13. [13]

    Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset, 2022

    Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset, 2022

  14. [14]

    Seedream 4.5.https://seed.bytedance.com/en/seedream4_5, 2026

    ByteDance Seed Team. Seedream 4.5.https://seed.bytedance.com/en/seedream4_5, 2026

  15. [15]

    Seedream 5.0 lite

    ByteDance Seed Team. Seedream 5.0 lite. https://seed.bytedance.com/en/blog/ deeper-thinking-more-accurate-generation-introducing-seedream-5-0-lite, 2026

  16. [16]

    Mm-iq: Benchmarking human-like abstraction and reasoning in multimodal models

    Huanqia Cai, Yijun Yang, and Winston Hu. Mm-iq: Benchmarking human-like abstraction and reasoning in multimodal models. arXiv preprint arXiv:2502.00698, 2025

  17. [17]

    Hidream-o1-image: A natively unified image generative foundation model with pixel-level unified transformer.arXiv preprint arXiv:2605.11061, 2026

    Qi Cai, Jingwen Chen, Chengmin Gao, Zijian Gong, Yehao Li, Tao Mei, Yingwei Pan, Yi Peng, Zhaofan Qiu, Ting Yao, Kai Yu, Yiheng Zhang, et al. Hidream-o1-image: A natively unified image generative foundation model with pixel-level unified transformer.arXiv preprint arXiv:2605.11061, 2026. 46

  18. [18]

    Rwku: Benchmarking real-world knowledge unlearning for large language models.Advances in Neural Information Processing Systems, 37:98213–98263, 2024

    Pengfei Cao, Chenhao Wang, Zhitao He, Hongbang Yuan, Jiachun Li, Yubo Chen, Kang Liu, Jun Zhao, et al. Rwku: Benchmarking real-world knowledge unlearning for large language models.Advances in Neural Information Processing Systems, 37:98213–98263, 2024

  19. [19]

    Hunyuanimage 3.0 technical report.arXiv preprint arXiv:2509.23951, 2025

    Siyu Cao, Hangting Chen, Peng Chen, Yiji Cheng, Yutao Cui, Xinchi Deng, Ying Dong, Kipper Gong, Tianpeng Gu, Xiusen Gu, et al. Hunyuanimage 3.0 technical report.arXiv preprint arXiv:2509.23951, 2025

  20. [20]

    Extracting training data from diffusion models

    Nicolas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramer, Borja Balle, Daphne Ippolito, and Eric Wallace. Extracting training data from diffusion models. In32nd USENIX security symposium (USENIX Security 23), pp. 5253–5270, 2023

  21. [21]

    Oneig-bench: Omni-dimensional nuanced evaluation for image generation, 2025

    Jingjing Chang, Yixiao Fang, Peng Xing, Shuhan Wu, Wei Cheng, Rui Wang, Xianfang Zeng, Gang Yu, and Hai-Bao Chen. Oneig-bench: Omni-dimensional nuanced evaluation for image generation, 2025. URL https://arxiv.org/abs/2506.07977

  22. [22]

    Lens: Rethinking training efficiency for foundational text-to-image models

    Dong Chen, Fangyun Wei, Ziyu Wan, Dongdong Chen, Jiawei Zhang, Jinjing Zhao, Sirui Zhang, Yang Yue, Zhiyang Liang, Baining Guo, et al. Lens: Rethinking training efficiency for foundational text-to-image models. arXiv preprint arXiv:2605.21573, 2026

  23. [23]

    Blip3o-next: Next frontier of native image generation, 2025

    Jiuhai Chen, Le Xue, Zhiyang Xu, Xichen Pan, Shusheng Yang, Can Qin, An Yan, Honglu Zhou, Zeyuan Chen, Lifu Huang, Tianyi Zhou, Junnan Li, Silvio Savarese, Caiming Xiong, and Ran Xu. Blip3o-next: Next frontier of native image generation, 2025. URLhttps://arxiv.org/abs/2510.15857

  24. [24]

    Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis, 2023

    Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis, 2023

  25. [25]

    Frugalgpt: How to use large language models while reducing cost and improving performance, 2023

    Lingjiao Chen, Matei Zaharia, and James Zou. Frugalgpt: How to use large language models while reducing cost and improving performance, 2023. URLhttps://arxiv.org/abs/2305.05176

  26. [26]

    Reproducible vision-language models meet concepts out of pre-training

    Ziliang Chen, Xin Huang, Xiaoxuan Fan, Keze Wang, Yuyu Zhou, Quanlong Guan, and Liang Lin. Reproducible vision-language models meet concepts out of pre-training. InProceedings of the Computer Vision and Pattern Recognition Conference, pp. 14701–14711, 2025

  27. [28]

    Goodhart’s law: its origins, meaning and implications for monetary policy

    K Alec Chrystal and Paul D Mizen. Goodhart’s law: its origins, meaning and implications for monetary policy. Central banking, monetary theory and practice: Essays in honour of Charles Goodhart, 1:221–243, 2003

  28. [29]

    Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, Rodolphe Jenatton, Lucas Beyer, Michael Tschannen, Anurag Arnab, Xiao Wang, Carlos Riquelme, Matthias Minderer, Joan Puigcerver, Utku Evci, Manoj Kumar, Sjoerd van Steenkiste, Gamaleldin F....

  29. [30]

    Imagenet: A large-scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009

  30. [31]

    Hybrid llm: Cost-efficient and quality-aware query routing

    Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Rühle, Laks Lakshmanan, and Ahmed H Awadallah. Hybrid llm: Cost-efficient and quality-aware query routing. InInternationalConference on Learning Representations, volume 2024, pp. 41348–41366, 2024

  31. [32]

    Cogview: Mastering text-to-image generation via transformers

    Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, and Jie Tang. Cogview: Mastering text-to-image generation via transformers. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (eds.),Advances in Neural Information Processing Systems, volume 34, pp. 19822–19835. Cu...

  32. [33]

    Documenting large webtext corpora: A case study on the colossal clean crawled corpus

    Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 conference on empirical methods in natural language processing, pp. 1286–1305, 2021

  33. [34]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. InInternational Conference on Learning Representations (ICLR), 2021. URL...

  34. [35]

    Textcrafter: Accurately rendering multiple texts in complex visual scenes.arXiv e-prints, pp

    Nikai Du, Zhennan Chen, Zhizhou Chen, Shan Gao, Xi Chen, Zhengkai Jiang, Jian Yang, and Ying Tai. Textcrafter: Accurately rendering multiple texts in complex visual scenes.arXiv e-prints, pp. arXiv–2503, 2025

  35. [36]

    Scaling rectified flow transformers for high-resolution image synthesis

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, and Robin Rombach. Scaling rectified flow transformers for high-resolution image synthesis. InProceedings of the 41st International Conference on Machine Learning, ICML...

  36. [37]

    Datacomp: in search of the next generation of multimodal datasets

    Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, Eyal Orgad, Rahim Entezari, Giannis Daras, Sarah Pratt, Vivek Ramanujan, Yonatan Bitton, Kalyani Marathe, Stephen Mussmann, Richard Vencu, Mehdi Cherti, Ranjay Krishna, Pang Wei Koh, Olga Saukh, Alexan...

  37. [38]

    Lumina-t2x: Scalable flow-based large diffusion transformer for flexible resolution generation

    Peng Gao, Le Zhuo, Dongyang Liu, Ruoyi Du, Xu Luo, Longtian Qiu, Yuhang Zhang, Rongjie Huang, Shijie Geng, Renrui Zhang, Junlin Xie, Wenqi Shao, Zhengkai Jiang, Tianshuo Yang, Weicai Ye, Tong He, Jingwen He, Junjun He, Yu Qiao, and Hongsheng Li. Lumina-t2x: Scalable flow-based large diffusion transformer for flexible resolution generation. InThe Thirteent...

  38. [39]

    Seedream 3.0 technical report, 2025

    Yu Gao, Lixue Gong, Qiushan Guo, Xiaoxia Hou, Zhichao Lai, Fanshi Li, Liang Li, Xiaochen Lian, Chao Liao, Liyang Liu, Wei Liu, Yichun Shi, Shiqi Sun, Yu Tian, Zhi Tian, Peng Wang, Rui Wang, Xuanda Wang, Xun Wang, Ye Wang, Guofeng Wu, Jie Wu, Xin Xia, Xuefeng Xiao, Zhonghua Zhai, Xinyu Zhang, Qi Zhang, Yuwei Zhang, Shijia Zhao, Jianchao Yang, and Weilin Hu...

  39. [40]

    Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities, 2025

    Google Gemini Team. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities, 2025. URLhttps://arxiv.org/abs/2507.06261

  40. [41]

    X-omni: Reinforcement learning makes discrete autoregressive image generative models great again

    Zigang Geng, Yibing Wang, Yeyao Ma, Chen Li, Yongming Rao, Shuyang Gu, Zhao Zhong, Qinglin Lu, Han Hu, Xiaosong Zhang, et al. X-omni: Reinforcement learning makes discrete autoregressive image generative models great again. arXiv preprint arXiv:2507.22058, 2025

  41. [42]

    Geneval: An object-focused framework for evaluating text-to-image alignment

    Dhruba Ghosh, Hannaneh Hajishirzi, and Ludwig Schmidt. Geneval: An object-focused framework for evaluating text-to-image alignment. Advances in Neural Information Processing Systems, 36:52132–52152, 2023

  42. [43]

    Introducing nano banana pro, 2025

    Google. Introducing nano banana pro, 2025. URL https://blog.google/innovation-and-ai/products/ nano-banana-pro/

  43. [44]

    Nano banana 2: Combining pro capabilities with lightning-fast speed, 2026

    Google. Nano banana 2: Combining pro capabilities with lightning-fast speed, 2026. URLhttps://blog.google/ innovation-and-ai/technology/ai/nano-banana-2/

  44. [45]

    Imagen 4.0.https://deepmind.google/models/imagen/, 2025

    Google DeepMind. Imagen 4.0.https://deepmind.google/models/imagen/, 2025

  45. [46]

    Gemini 3.1 pro - model card

    Google DeepMind. Gemini 3.1 pro - model card. https://deepmind.google/models/model-cards/ gemini-3-1-pro/, February 2026. Accessed: 2026-05-05

  46. [47]

    Accurate, large minibatch sgd: Training imagenet in 1 hour, 2018

    Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. Accurate, large minibatch sgd: Training imagenet in 1 hour, 2018. URL https://arxiv.org/abs/1706.02677. 48

  47. [48]

    Accelerate: Training and inference at scale made simple, efficient and adaptable

    Sylvain Gugger, Lysandre Debut, Thomas Wolf, Philipp Schmid, Zachary Mueller, Sourab Mangrulkar, Marc Sun, and Benjamin Bossan. Accelerate: Training and inference at scale made simple, efficient and adaptable. https://github.com/huggingface/accelerate, 2022

  48. [49]

    R-bench: Graduate-level multi-disciplinary benchmarks for llm & mllm complex reasoning evaluation.arXiv preprint arXiv:2505.02018, 2025

    Meng-Hao Guo, Jiajun Xu, Yi Zhang, Jiaxi Song, Haoyang Peng, Yi-Xuan Deng, Xinzhi Dong, Kiyohiro Nakayama, Zhengyang Geng, Chen Wang, et al. R-bench: Graduate-level multi-disciplinary benchmarks for llm & mllm complex reasoning evaluation.arXiv preprint arXiv:2505.02018, 2025

  49. [50]

    Shampoo: Preconditioned stochastic tensor optimization, 2018

    Vineet Gupta, Tomer Koren, and Yoram Singer. Shampoo: Preconditioned stochastic tensor optimization, 2018. URLhttps://arxiv.org/abs/1802.09568

  50. [51]

    Optimizing prompts for text-to-image generation.Advances in Neural Information Processing Systems, 36:66923–66939, 2023

    Yaru Hao, Zewen Chi, Li Dong, and Furu Wei. Optimizing prompts for text-to-image generation.Advances in Neural Information Processing Systems, 36:66923–66939, 2023

  51. [52]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016

  52. [53]

    Gems: Agent-native multimodal generation with memory and skills.arXiv preprint arXiv:2603.28088, 2026

    Zefeng He, Siyuan Huang, Xiaoye Qu, Yafu Li, Tong Zhu, Yu Cheng, and Yang Yang. Gems: Agent-native multimodal generation with memory and skills.arXiv preprint arXiv:2603.28088, 2026

  53. [54]

    Clipscore: A reference-free evaluation metric for image captioning

    Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation metric for image captioning. InProceedings of the 2021 conference on empirical methods in natural language processing, pp. 7514–7528, 2021

  54. [55]

    Classifier-free diffusion guidance

    Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. InNeurIPS2021WorkshoponDeepGenerative Models and Downstream Applications, 2021. URLhttps://openreview.net/forum?id=qw8AKxfYbI

  55. [56]

    Denoising diffusion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. InProceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20, Red Hook, NY, USA, 2020. Curran Associates Inc. ISBN 9781713829546

  56. [57]

    Ella: Equip diffusion models with llm for enhanced semantic alignment.arXiv preprint arXiv:2403.05135, 2024

    Xiwei Hu, Rui Wang, Yixiao Fang, Bin Fu, Pei Cheng, and Gang Yu. Ella: Equip diffusion models with llm for enhanced semantic alignment.arXiv preprint arXiv:2403.05135, 2024

  57. [58]

    Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering

    Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Ostendorf, Ranjay Krishna, and Noah A Smith. Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering. InProceedings of the IEEE/CVF International Conference on Computer Vision, pp. 20406–20417, 2023

  58. [59]

    T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.Advancesin Neural Information Processing Systems, 36: 78723–78747, 2023

    Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu. T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.Advancesin Neural Information Processing Systems, 36: 78723–78747, 2023

  59. [60]

    Ape: Agentic prompt enhancer for image generation and editing.arXiv preprint arXiv:2606.00204, 2026

    Zijian Huang, Jay Zhangjie Wu, Zian Wang, Tianshi Cao, Jiasi Chen, Sanja Fidler, Huan Ling, and Xuanchi Ren. Ape: Agentic prompt enhancer for image generation and editing.arXiv preprint arXiv:2606.00204, 2026

  60. [61]

    Scaling up visual and vision-language representation learning with noisy text supervision

    Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International conference on machine learning, pp. 4904–4916. PMLR, 2021

  61. [62]

    Geditbench v2: A human-aligned benchmark for general image editing.arXiv preprint arXiv:2603.28547, 2026

    Zhangqi Jiang, Zheng Sun, Xianfang Zeng, Yufeng Yang, Xuanyang Zhang, Yongliang Wu, Wei Cheng, Gang Yu, Xu Yang, and Bihan Wen. Geditbench v2: A human-aligned benchmark for general image editing.arXiv preprint arXiv:2603.28547, 2026

  62. [63]

    Muon: An optimizer for hidden layers in neural networks, 2024

    Keller Jordan, Yuchen Jin, Vlado Boza, Jiacheng You, Franz Cesista, Laker Newhouse, and Jeremy Bernstein. Muon: An optimizer for hidden layers in neural networks, 2024. URLhttps://kellerjordan.github.io/posts/ muon/

  63. [64]

    Scaling laws for neural language models

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020

  64. [65]

    Elucidating the design space of diffusion-based generative models

    Tero Karras, Miika Aittala, Samuli Laine, and Timo Aila. Elucidating the design space of diffusion-based generative models. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22, Red Hook, NY, USA, 2022. Curran Associates Inc. ISBN 9781713871088. 49

  65. [66]

    Analyzing and improving the training dynamics of diffusion models, 2024

    Tero Karras, Miika Aittala, Jaakko Lehtinen, Janne Hellsten, Timo Aila, and Samuli Laine. Analyzing and improving the training dynamics of diffusion models, 2024. URLhttps://arxiv.org/abs/2312.02696

  66. [67]

    Auto-encoding variational bayes.arXiv preprint arXiv:1312.6114, 2013

    Diederik P Kingma and Max Welling. Auto-encoding variational bayes.arXiv preprint arXiv:1312.6114, 2013

  67. [68]

    Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advancesin neural information processing systems, 36:36652–36663, 2023

    Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advancesin neural information processing systems, 36:36652–36663, 2023

  68. [69]

    Imagenet classification with deep convolutional neural networks

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012

  69. [70]

    Kling-image-2.1.https://kling.ai/explore/kling_2.1_api, 2025

    Kuaishou Technology. Kling-image-2.1.https://kling.ai/explore/kling_2.1_api, 2025

  70. [71]

    Improved precision and recall metric for assessing generative models.Advances in neural information processing systems, 32, 2019

    Tuomas Kynkäänniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, and Timo Aila. Improved precision and recall metric for assessing generative models.Advances in neural information processing systems, 32, 2019

  71. [72]

    Krea 2.https://www.krea.ai/blog/krea-2-technical-report, 2026

    Sangwu Lee, Erwann Millon, Le Zhuo, Matthew Newton, Andrei Filatov, Naga Sai Abhinay Devarinti, Dazhi Zhong, Avram Djordjevic, Gabriel Menezes, Will Beddow, Titus Ebbecke, Mihai Petrescu, Owen Fahey, Gian Saß, Felix Gil, and Victor Perez. Krea 2.https://www.krea.ai/blog/krea-2-technical-report, 2026

  72. [73]

    Qwen-image-bench: From generation to creation in text-to-image evaluation,

    Niantong Li, Guangzheng Hu, Weixu Qiao, Ying Ba, Qichen Hong, Shijun Shen, Jinlin Wang, Fan Zhou, Jianye Kang, Xin Shang, Ziyi He, Wei Wang, Dalin Li, Jiahao Li, Jie Zhang, Kaiyuan Gao, Kun Yan, Lihan Jiang, Ningyuan Tang, Shengming Yin, Tianhe Wu, Xiao Xu, Xiaoyue Chen, Yuxiang Chen, Yan Shu, Yanran Zhang, Yilei Chen, Yixian Xu, Zekai Zhang, Zhendong Wan...

  73. [74]

    Omnicorpus: A unified multimodal corpus of 10 billion-level images interleaved with text

    Qingyun Li, Zhe Chen, Weiyun Wang, Wenhai Wang, Shenglong Ye, Zhenjiang Jin, Guanzhou Chen, Yinan He, Zhangwei Gao, Erfei Cui, Jiashuo Yu, Hao Tian, Jiasheng Zhou, Chao Xu, Bin Wang, Xingjian Wei, Wei Li, Wenjian Zhang, Bo Zhang, Pinlong Cai, Licheng Wen, Xiangchao Yan, Pei Chu, Yi Wang, Min Dou, Changyao Tian, Xizhou Zhu, Lewei Lu, Yushi Chen, Junjun He,...

  74. [75]

    Reflect-dit: Inference-time scaling for text-to-image diffusion transformers via in-context reflection

    Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Arsh Koneru, Yusuke Kato, Kazuki Kozuka, and Aditya Grover. Reflect-dit: Inference-time scaling for text-to-image diffusion transformers via in-context reflection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 15657–15668, 2025

  75. [76]

    Holistic evaluation of language models.arXiv preprint arXiv:2211.09110, 2022

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. Holistic evaluation of language models.arXiv preprint arXiv:2211.09110, 2022

  76. [77]

    Scaling laws for diffusion transformers.arXiv preprint arXiv:2410.08184, 2024

    Zhengyang Liang, Hao He, Ceyuan Yang, and Bo Dai. Scaling laws for diffusion transformers.arXiv preprint arXiv:2410.08184, 2024

  77. [78]

    Aegis: Exploring the limit of world knowledge capabilities for unified mulitmodal models.arXiv preprint arXiv:2601.00561, 2026

    Jintao Lin, Bowen Dong, Weikang Shi, Chenyang Lei, Suiyun Zhang, Rui Liu, and Xihui Liu. Aegis: Exploring the limit of world knowledge capabilities for unified mulitmodal models.arXiv preprint arXiv:2601.00561, 2026

  78. [79]

    Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling, 2023. URLhttps://arxiv.org/abs/2210.02747

  79. [80]

    Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

  80. [81]

    Ernie-image technical report.arXiv preprint arXiv:2605.25347, 2026

    Jiaxiang Liu, Zhida Feng, Pengyu Zou, Zhenyu Qian, Tianrui Zhu, Jun Xia, Yuehu Dong, Yanzheng Lin, Honglin Xiong, Anqi Chen, et al. Ernie-image technical report.arXiv preprint arXiv:2605.25347, 2026

Showing first 80 references.