REVIEW 4 major objections 5 minor 153 references
Boogu-Image-0.1 argues that upgrading the understanding side of a text-to-image system—encoder, prompt rewriting, captions, routing—can make a model trained on 208.62M images at roughly $400K competitive with far costlier open- and closed-s
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 06:07 UTC pith:2JDMA2ZH
load-bearing objection Genuinely useful engineering report with honest ablations; flagship SOTA claim rests on an in-house benchmark needing external validation. the 4 major comments →
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On its own terms, the paper establishes that understanding can be engineered as a first-class, decoupled component of a generation system rather than left implicit in a large generator. It decomposes understanding into an instruction encoder that acts as a lossy sensor, an agentic prompt rewriter that translates intent without embellishing clear prompts, captions that name what users will ask for (including visual flaws), and a difficulty-aware router that dispatches easy requests to fast models. The empirical core is a series of controlled comparisons: encoder scale alone moves GenEval from 0.60 to 0.65 with no saturation; the same 47M-image syllabus outperforms 187M of open-source data; th
What carries the argument
The load-bearing mechanism is a decoupled understanding stack placed upstream of generation. A mid-sized frozen vision-language model serves as the instruction encoder, treated as a sensor whose capacity upper-bounds what the diffusion transformer can receive, while a separate reasoning-capable agentic rewriter and a per-aspect captioning pipeline supply the conditioning and supervision. Completing the stack is a router that classifies request difficulty and selects fast versus heavy model variants, turning inference time into a dial on a quality-latency curve. Two supporting techniques carry much of the training-efficiency claim: a syllabus-guided data curriculum with explicit defect captio
Load-bearing premise
The headline top-open-source-model ranking depends on Boogu Arena, an in-house Elo benchmark whose agreement with public human preference was measured on only five models; if that arena does not reflect real user preference, the central competitive claim is unsupported.
What would settle it
Run a public, independent blind pairwise comparison, for example on a held-out bilingual prompt set, between Boogu-Image-0.1-Turbo-Thinking and the leading open baselines it claims to beat, and check whether its Elo still leads; since the weights are released, this is directly executable.
If this is right
- Open-source teams can budget for roughly 200M unique images and about $400K of compute and still reach the top of the open tier, provided understanding-side components are treated as part of the model.
- Evaluation that reports quality without latency is incomplete: the same Boogu model occupies many points on the quality-time curve, and thinking and routing variants make that curve explicit.
- Capabilities that are not named in captions cannot be recovered at inference time, so captioning should be planned per anticipated user demand rather than delegated to a single vision-language model.
- Saturated academic benchmarks such as GenEval and DPG-Bench no longer track human preference, so the paper's results on its arena and on a post-freeze open benchmark carry the evidential weight.
- Skills embedded in the rewriter, such as counting, infographic layout, reasoning, NSFW handling, and scene text, transfer into measurable generation gains in those same dimensions.
Where Pith is reading between the lines
- If the roughly 300-exposure glyph result generalizes, it predicts a cheap and testable recipe for any new writing system or token set: enumerate the vocabulary and floor each unit's exposure rather than scaling raw data.
- The reported inversion on the editing benchmark, where a closed system scores lower yet wins in the authors' own human evaluation, suggests that model-based editing metrics systematically compress quality gaps; a public human editing arena could settle which ranking is real.
- The finding that heavy aesthetic reinforcement learning narrows output distributions implies a policy tension: optimizing average appeal can erase off-mainstream demographics and styles, which matters for fairness and access as much as for diversity.
- Releasing weights with full recipes makes the understanding-first budget claim directly falsifiable and re-runnable, which is the main reason the cost figure can be checked rather than taken on faith.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Boogu-Image-0.1, an open-source unified text-to-image and image-editing model family (Base, Turbo, Edit, Edit-Turbo), and argues that strengthening the understanding component — via a stronger frozen multimodal encoder (Qwen3-VL-8B), an agentic prompt rewriter, a curated data syllabus, and agentic inference-time scaling — substantially improves generation quality under a minimal compute budget (208.62M unique images, ~$400K theoretical training cost). The evaluation relies on an in-house blind pairwise benchmark (Boogu Arena), on external benchmarks Qwen-Image-Bench and LongText-Bench, and on ImgEdit-Bench for editing, claiming top-tier open-source status and performance approaching closed-source systems.
Significance. If the central claims hold, the paper offers two useful contributions: (i) a systematic demonstration, with ablations, that instruction-encoder scale, curated data syllabus, and agentic prompt rewriting each yield measurable gains under a fixed budget, and (ii) a practical open release of weights, code, and training recipes. The ablations in Tables 7–8 and Figure 18 are genuine and non-circular, comparing actual checkpoints and data configurations. The bottleneck is the headline positioning: the claim of "approaching closed-source systems" and of open-source leadership on human preference rests almost entirely on an in-house benchmark whose external validation is thin, and the paper itself documents an inversion between its automated editing benchmark and its own human evaluation. The external Qwen-Image-Bench results support open-source leadership only among the subset of models evaluated and do not support the closed-source gap claim.
major comments (4)
- [§2.2.1, Figure 7] The Boogu Arena Elo ranking is the primary evidence for the headline "approaching closed-source" claim, but its external validation is limited to five models, none of which are Boogu models. With n=5, a Pearson r of 0.986 and Spearman 1.0 are consistent with a linear trend but do not establish that Boogu's absolute Elo position transfers to the LMArena scale; the Boogu points are extrapolations. The paper reports no confidence intervals on Elo, no inter-annotator agreement, and no independent annotation protocol, and the prompts are only promised for future release. The paper itself (Sec. 2.3.1) shows that its in-house preference instrument can invert the ranking of two models relative to human evaluation, so Boogu Arena cannot currently carry the closed-source gap claim.
- [§2.3.1, Table 6] The text states that Boogu-Image-0.1-Edit-Thinking "attains the best overall score" on ImgEdit-Bench, but the immediately following Discussion and Limitations says that in the authors' own human evaluations Nano-Banana-Pro outperforms Boogu despite scoring lower (4.37 vs 4.64), and the table is labeled "for reference only." This is not a minor caveat: the abstract and introduction use "matches or surpasses" across standard benchmarks without this qualification, and the internal inversion means the ImgEdit ranking should not be presented as a headline result. The claim needs to be recast or the human evaluation data reported.
- [§2.2.2, Tables 1–2; §2.2.3, Table 3] The open-source leadership claim on Qwen-Image-Bench and LongText-Bench is limited by the evaluated set: the text says "due to time constraints, the evaluation does not yet cover all available open-source baselines," and no inference-time budget is reported for any model. Since the paper argues (Sec. 3.2.1) that quality must be reported jointly with latency, the thinking variants, which use agentic rewriting and reflection, should be compared with closed-source systems under a matched or explicitly reported inference budget; otherwise the 'approaching closed-source' portion of the claim is not supported by these tables.
- [§3.2.3 (Dynamic time shifting) and Figure 32] The rectified dynamic time-shifting scheme introduces a free parameter ncap (e.g., 4096), and its benefit is illustrated by quantile plots rather than by a controlled training ablation showing that the rectified scheme improves final generation metrics (e.g., FID or benchmark scores) relative to standard dynamic shifting at 2K. The claim that this is a decisive detail for the Boogu pipeline is plausible but currently under-supported; either add an ablation or soften the presentation.
minor comments (5)
- [Figure 1] The right panel lists 'Boogu-Image-0.1-Turbo-PE' but the model family in the abstract and Section 2 includes Base, Turbo, Edit, Edit-Turbo; PE is not defined in the caption or text. Please clarify or rename.
- [Tables 1–6] Model names are inconsistent: 'Nano-Banana-Pro' vs 'NanoBanana-Pro', and 'Qwen-Image-2.0-2026-03-03' vs 'Qwen-Image-2.0'. Also, some tables say 'Our model series achieves...' while the model is highlighted in red; ensure the caption convention is uniform.
- [§2.2.4] The decision to report GenEval and DPG-Bench 'for reference only' is reasonable, but the same standards should apply to ImgEdit-Bench, which is also based on VLM scoring and acknowledged to invert human preference. Consider adding a unified statement about which benchmarks the authors regard as primary and why.
- [§4] The limitations paragraph is candid but does not mention the absence of an external human-preference validation for Boogu Arena; this is the most load-bearing limitation for the abstract's claim and should be stated there.
- [Abstract] The phrase 'the base model's theoretical training cost is only approximately $400K' is used without defining what 'theoretical' includes (hardware, electricity, ablations, data curation). Please provide a cost breakdown or soften the claim.
Circularity Check
No significant circularity: central claims rest on independent ablations and external benchmarks; Boogu Arena is a validity concern, not a circular reduction.
full rationale
The paper's load-bearing derivations are direct empirical ablations over actual checkpoints and data configurations: Table 7 varies the frozen instruction encoder and measures GenEval; Figure 18 scales the rewriter backbone and reports a defined advantage score; Table 8 compares training on 187M open-source data vs. the Boogu Syllabus on Qwen-Image-Bench and LongText-Bench; Figures 27-28 and Table 9 measure glyph-exposure thresholds. In each case the manipulated variable is not defined in terms of the measured outcome, so no self-definitional reduction is present. The supporting benchmarks (Qwen-Image-Bench, LongText-Bench, GenEval, DPG-Bench) are external to the paper. The Boogu Arena leaderboard is built in-house, but its Elo scores are not fitted to Boogu outputs, and the LMArena agreement check (Figure 7), while based on only five non-Boogu models, is a calibration exercise rather than a definitional reduction. The only apparent self-citation ([78], used as an example in the open-challenges discussion) is not load-bearing. Benchmark-validity concerns about Boogu Arena are legitimate evaluation risks, but they are not circularity under the definitions used here.
Axiom & Free-Parameter Ledger
free parameters (1)
- ncap (dynamic time-shift clamp) =
4096
axioms (4)
- domain assumption The Boogu Arena Elo computed from in-house blind pairwise votes is a faithful proxy for real-world human preference (correlation with LMArena r=0.986 on five anchor models).
- domain assumption VLM/automated benchmarks (Qwen-Image-Bench, LongText-Bench, ImgEdit-Bench) are meaningful for the claims despite acknowledged saturation and their own ImgEdit inversion.
- domain assumption A per-exposure threshold (300 per glyph; ~4,500 per identity) measured on a small subset generalizes to the full training distribution.
- standard math Standard flow-matching/diffusion and DiT formulations are correct as used.
read the original abstract
We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual (Chinese-English) text rendering. Closed-source multimodal systems like Nano-Banana-Pro and GPT-Image-2 achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. In this work, we demonstrate that strengthening the understanding capability of the system, through a stronger multimodal encoder, agentic prompt rewriting, and related techniques, together with improvements in data quality, training pipelines, and agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets. Comprehensive evaluations show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks, and achieves results approaching leading closed-source systems. Notably, this is accomplished with only 208.62 million unique images. The base model's theoretical training cost is only approximately \$400K. We share practical discussions that we believe are valuable to the broader research community, and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. Our code is available here: https://github.com/Boogu-Project/Boogu-Image.
Reference graph
Works this paper leans on
-
[1]
Arena ai: The official ai ranking & llm leaderboard.https://arena.ai/, 2026
2026
-
[2]
Ideogram 4.https://ideogram.ai/blog/ideogram-4.0/, 2026
Ideogram AI. Ideogram 4.https://ideogram.ai/blog/ideogram-4.0/, 2026
2026
-
[3]
Qwen3-vl technical report, 2025
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan ...
Pith/arXiv arXiv 2025
-
[4]
Qwen2.5-vl technical report, 2025
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report, 2...
Pith/arXiv arXiv 2025
-
[5]
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
BIG bench authors. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URLhttps://openreview.net/ forum?id=uyTL5Bvosj
2023
-
[6]
Improving image generation with better captions.Computer Science
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al. Improving image generation with better captions.Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3):8, 2023
2023
-
[7]
Announcing Black Forest Labs: Black forest labs’ inaugural model suite, FLUX.1, 2024
Black Forest Labs. Announcing Black Forest Labs: Black forest labs’ inaugural model suite, FLUX.1, 2024. URL https://blackforestlabs.ai/announcing-black-forest-labs/. Accessed: 2026-04-25
2024
-
[8]
FLUX.2: Analyzing and enhancing the latent space of FLUX – representation comparison, 11
Black Forest Labs. FLUX.2: Analyzing and enhancing the latent space of FLUX – representation comparison, 11
-
[9]
Flux.2: Production-grade ai image generation and editing model with 4mp photorealistic output and multi-reference control.https://bfl.ai/models/flux-2, 2025
Black Forest Labs. Flux.2: Production-grade ai image generation and editing model with 4mp photorealistic output and multi-reference control.https://bfl.ai/models/flux-2, 2025
2025
-
[10]
Large scale GAN training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale GAN training for high fidelity natural image synthesis. In International Conference on Learning Representations, 2019. URL https://openreview.net/ forum?id=B1xsqj09Fm
2019
-
[11]
Instructpix2pix: Learning to follow image editing instructions
Tim Brooks, Aleksander Holynski, and Alexei A Efros. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 18392–18402, 2023
2023
-
[12]
Video generation models as world simulators.OpenAI Blog, 1(8):1, 2024
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Leo Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, et al. Video generation models as world simulators.OpenAI Blog, 1(8):1, 2024
2024
-
[13]
Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset, 2022
Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset, 2022
2022
-
[14]
Seedream 4.5.https://seed.bytedance.com/en/seedream4_5, 2026
ByteDance Seed Team. Seedream 4.5.https://seed.bytedance.com/en/seedream4_5, 2026
2026
-
[15]
Seedream 5.0 lite
ByteDance Seed Team. Seedream 5.0 lite. https://seed.bytedance.com/en/blog/ deeper-thinking-more-accurate-generation-introducing-seedream-5-0-lite, 2026
2026
-
[16]
Mm-iq: Benchmarking human-like abstraction and reasoning in multimodal models
Huanqia Cai, Yijun Yang, and Winston Hu. Mm-iq: Benchmarking human-like abstraction and reasoning in multimodal models. arXiv preprint arXiv:2502.00698, 2025
Pith/arXiv arXiv 2025
-
[17]
Qi Cai, Jingwen Chen, Chengmin Gao, Zijian Gong, Yehao Li, Tao Mei, Yingwei Pan, Yi Peng, Zhaofan Qiu, Ting Yao, Kai Yu, Yiheng Zhang, et al. Hidream-o1-image: A natively unified image generative foundation model with pixel-level unified transformer.arXiv preprint arXiv:2605.11061, 2026. 46
Pith/arXiv arXiv 2026
-
[18]
Rwku: Benchmarking real-world knowledge unlearning for large language models.Advances in Neural Information Processing Systems, 37:98213–98263, 2024
Pengfei Cao, Chenhao Wang, Zhitao He, Hongbang Yuan, Jiachun Li, Yubo Chen, Kang Liu, Jun Zhao, et al. Rwku: Benchmarking real-world knowledge unlearning for large language models.Advances in Neural Information Processing Systems, 37:98213–98263, 2024
2024
-
[19]
Hunyuanimage 3.0 technical report.arXiv preprint arXiv:2509.23951, 2025
Siyu Cao, Hangting Chen, Peng Chen, Yiji Cheng, Yutao Cui, Xinchi Deng, Ying Dong, Kipper Gong, Tianpeng Gu, Xiusen Gu, et al. Hunyuanimage 3.0 technical report.arXiv preprint arXiv:2509.23951, 2025
Pith/arXiv arXiv 2025
-
[20]
Extracting training data from diffusion models
Nicolas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramer, Borja Balle, Daphne Ippolito, and Eric Wallace. Extracting training data from diffusion models. In32nd USENIX security symposium (USENIX Security 23), pp. 5253–5270, 2023
2023
-
[21]
Oneig-bench: Omni-dimensional nuanced evaluation for image generation, 2025
Jingjing Chang, Yixiao Fang, Peng Xing, Shuhan Wu, Wei Cheng, Rui Wang, Xianfang Zeng, Gang Yu, and Hai-Bao Chen. Oneig-bench: Omni-dimensional nuanced evaluation for image generation, 2025. URL https://arxiv.org/abs/2506.07977
Pith/arXiv arXiv 2025
-
[22]
Lens: Rethinking training efficiency for foundational text-to-image models
Dong Chen, Fangyun Wei, Ziyu Wan, Dongdong Chen, Jiawei Zhang, Jinjing Zhao, Sirui Zhang, Yang Yue, Zhiyang Liang, Baining Guo, et al. Lens: Rethinking training efficiency for foundational text-to-image models. arXiv preprint arXiv:2605.21573, 2026
Pith/arXiv arXiv 2026
-
[23]
Blip3o-next: Next frontier of native image generation, 2025
Jiuhai Chen, Le Xue, Zhiyang Xu, Xichen Pan, Shusheng Yang, Can Qin, An Yan, Honglu Zhou, Zeyuan Chen, Lifu Huang, Tianyi Zhou, Junnan Li, Silvio Savarese, Caiming Xiong, and Ran Xu. Blip3o-next: Next frontier of native image generation, 2025. URLhttps://arxiv.org/abs/2510.15857
arXiv 2025
-
[24]
Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis, 2023
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis, 2023
2023
-
[25]
Frugalgpt: How to use large language models while reducing cost and improving performance, 2023
Lingjiao Chen, Matei Zaharia, and James Zou. Frugalgpt: How to use large language models while reducing cost and improving performance, 2023. URLhttps://arxiv.org/abs/2305.05176
Pith/arXiv arXiv 2023
-
[26]
Reproducible vision-language models meet concepts out of pre-training
Ziliang Chen, Xin Huang, Xiaoxuan Fan, Keze Wang, Yuyu Zhou, Quanlong Guan, and Liang Lin. Reproducible vision-language models meet concepts out of pre-training. InProceedings of the Computer Vision and Pattern Recognition Conference, pp. 14701–14711, 2025
2025
-
[28]
Goodhart’s law: its origins, meaning and implications for monetary policy
K Alec Chrystal and Paul D Mizen. Goodhart’s law: its origins, meaning and implications for monetary policy. Central banking, monetary theory and practice: Essays in honour of Charles Goodhart, 1:221–243, 2003
2003
-
[29]
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, Rodolphe Jenatton, Lucas Beyer, Michael Tschannen, Anurag Arnab, Xiao Wang, Carlos Riquelme, Matthias Minderer, Joan Puigcerver, Utku Evci, Manoj Kumar, Sjoerd van Steenkiste, Gamaleldin F....
-
[30]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009
2009
-
[31]
Hybrid llm: Cost-efficient and quality-aware query routing
Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Rühle, Laks Lakshmanan, and Ahmed H Awadallah. Hybrid llm: Cost-efficient and quality-aware query routing. InInternationalConference on Learning Representations, volume 2024, pp. 41348–41366, 2024
2024
-
[32]
Cogview: Mastering text-to-image generation via transformers
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, and Jie Tang. Cogview: Mastering text-to-image generation via transformers. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (eds.),Advances in Neural Information Processing Systems, volume 34, pp. 19822–19835. Cu...
2021
-
[33]
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 conference on empirical methods in natural language processing, pp. 1286–1305, 2021
2021
-
[34]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. InInternational Conference on Learning Representations (ICLR), 2021. URL...
2021
-
[35]
Textcrafter: Accurately rendering multiple texts in complex visual scenes.arXiv e-prints, pp
Nikai Du, Zhennan Chen, Zhizhou Chen, Shan Gao, Xi Chen, Zhengkai Jiang, Jian Yang, and Ying Tai. Textcrafter: Accurately rendering multiple texts in complex visual scenes.arXiv e-prints, pp. arXiv–2503, 2025
2025
-
[36]
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, and Robin Rombach. Scaling rectified flow transformers for high-resolution image synthesis. InProceedings of the 41st International Conference on Machine Learning, ICML...
2024
-
[37]
Datacomp: in search of the next generation of multimodal datasets
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, Eyal Orgad, Rahim Entezari, Giannis Daras, Sarah Pratt, Vivek Ramanujan, Yonatan Bitton, Kalyani Marathe, Stephen Mussmann, Richard Vencu, Mehdi Cherti, Ranjay Krishna, Pang Wei Koh, Olga Saukh, Alexan...
2023
-
[38]
Lumina-t2x: Scalable flow-based large diffusion transformer for flexible resolution generation
Peng Gao, Le Zhuo, Dongyang Liu, Ruoyi Du, Xu Luo, Longtian Qiu, Yuhang Zhang, Rongjie Huang, Shijie Geng, Renrui Zhang, Junlin Xie, Wenqi Shao, Zhengkai Jiang, Tianshuo Yang, Weicai Ye, Tong He, Jingwen He, Junjun He, Yu Qiao, and Hongsheng Li. Lumina-t2x: Scalable flow-based large diffusion transformer for flexible resolution generation. InThe Thirteent...
2025
-
[39]
Seedream 3.0 technical report, 2025
Yu Gao, Lixue Gong, Qiushan Guo, Xiaoxia Hou, Zhichao Lai, Fanshi Li, Liang Li, Xiaochen Lian, Chao Liao, Liyang Liu, Wei Liu, Yichun Shi, Shiqi Sun, Yu Tian, Zhi Tian, Peng Wang, Rui Wang, Xuanda Wang, Xun Wang, Ye Wang, Guofeng Wu, Jie Wu, Xin Xia, Xuefeng Xiao, Zhonghua Zhai, Xinyu Zhang, Qi Zhang, Yuwei Zhang, Shijia Zhao, Jianchao Yang, and Weilin Hu...
Pith/arXiv arXiv 2025
-
[40]
Google Gemini Team. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities, 2025. URLhttps://arxiv.org/abs/2507.06261
Pith/arXiv arXiv 2025
-
[41]
X-omni: Reinforcement learning makes discrete autoregressive image generative models great again
Zigang Geng, Yibing Wang, Yeyao Ma, Chen Li, Yongming Rao, Shuyang Gu, Zhao Zhong, Qinglin Lu, Han Hu, Xiaosong Zhang, et al. X-omni: Reinforcement learning makes discrete autoregressive image generative models great again. arXiv preprint arXiv:2507.22058, 2025
Pith/arXiv arXiv 2025
-
[42]
Geneval: An object-focused framework for evaluating text-to-image alignment
Dhruba Ghosh, Hannaneh Hajishirzi, and Ludwig Schmidt. Geneval: An object-focused framework for evaluating text-to-image alignment. Advances in Neural Information Processing Systems, 36:52132–52152, 2023
2023
-
[43]
Introducing nano banana pro, 2025
Google. Introducing nano banana pro, 2025. URL https://blog.google/innovation-and-ai/products/ nano-banana-pro/
2025
-
[44]
Nano banana 2: Combining pro capabilities with lightning-fast speed, 2026
Google. Nano banana 2: Combining pro capabilities with lightning-fast speed, 2026. URLhttps://blog.google/ innovation-and-ai/technology/ai/nano-banana-2/
2026
-
[45]
Imagen 4.0.https://deepmind.google/models/imagen/, 2025
Google DeepMind. Imagen 4.0.https://deepmind.google/models/imagen/, 2025
2025
-
[46]
Gemini 3.1 pro - model card
Google DeepMind. Gemini 3.1 pro - model card. https://deepmind.google/models/model-cards/ gemini-3-1-pro/, February 2026. Accessed: 2026-05-05
2026
-
[47]
Accurate, large minibatch sgd: Training imagenet in 1 hour, 2018
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. Accurate, large minibatch sgd: Training imagenet in 1 hour, 2018. URL https://arxiv.org/abs/1706.02677. 48
Pith/arXiv arXiv 2018
-
[48]
Accelerate: Training and inference at scale made simple, efficient and adaptable
Sylvain Gugger, Lysandre Debut, Thomas Wolf, Philipp Schmid, Zachary Mueller, Sourab Mangrulkar, Marc Sun, and Benjamin Bossan. Accelerate: Training and inference at scale made simple, efficient and adaptable. https://github.com/huggingface/accelerate, 2022
2022
-
[49]
Meng-Hao Guo, Jiajun Xu, Yi Zhang, Jiaxi Song, Haoyang Peng, Yi-Xuan Deng, Xinzhi Dong, Kiyohiro Nakayama, Zhengyang Geng, Chen Wang, et al. R-bench: Graduate-level multi-disciplinary benchmarks for llm & mllm complex reasoning evaluation.arXiv preprint arXiv:2505.02018, 2025
Pith/arXiv arXiv 2025
-
[50]
Shampoo: Preconditioned stochastic tensor optimization, 2018
Vineet Gupta, Tomer Koren, and Yoram Singer. Shampoo: Preconditioned stochastic tensor optimization, 2018. URLhttps://arxiv.org/abs/1802.09568
Pith/arXiv arXiv 2018
-
[51]
Optimizing prompts for text-to-image generation.Advances in Neural Information Processing Systems, 36:66923–66939, 2023
Yaru Hao, Zewen Chi, Li Dong, and Furu Wei. Optimizing prompts for text-to-image generation.Advances in Neural Information Processing Systems, 36:66923–66939, 2023
2023
-
[52]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016
2016
-
[53]
Zefeng He, Siyuan Huang, Xiaoye Qu, Yafu Li, Tong Zhu, Yu Cheng, and Yang Yang. Gems: Agent-native multimodal generation with memory and skills.arXiv preprint arXiv:2603.28088, 2026
arXiv 2026
-
[54]
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation metric for image captioning. InProceedings of the 2021 conference on empirical methods in natural language processing, pp. 7514–7528, 2021
2021
-
[55]
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. InNeurIPS2021WorkshoponDeepGenerative Models and Downstream Applications, 2021. URLhttps://openreview.net/forum?id=qw8AKxfYbI
2021
-
[56]
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. InProceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20, Red Hook, NY, USA, 2020. Curran Associates Inc. ISBN 9781713829546
2020
-
[57]
Xiwei Hu, Rui Wang, Yixiao Fang, Bin Fu, Pei Cheng, and Gang Yu. Ella: Equip diffusion models with llm for enhanced semantic alignment.arXiv preprint arXiv:2403.05135, 2024
Pith/arXiv arXiv 2024
-
[58]
Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering
Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Ostendorf, Ranjay Krishna, and Noah A Smith. Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering. InProceedings of the IEEE/CVF International Conference on Computer Vision, pp. 20406–20417, 2023
2023
-
[59]
T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.Advancesin Neural Information Processing Systems, 36: 78723–78747, 2023
Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu. T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.Advancesin Neural Information Processing Systems, 36: 78723–78747, 2023
2023
-
[60]
Ape: Agentic prompt enhancer for image generation and editing.arXiv preprint arXiv:2606.00204, 2026
Zijian Huang, Jay Zhangjie Wu, Zian Wang, Tianshi Cao, Jiasi Chen, Sanja Fidler, Huan Ling, and Xuanchi Ren. Ape: Agentic prompt enhancer for image generation and editing.arXiv preprint arXiv:2606.00204, 2026
Pith/arXiv arXiv 2026
-
[61]
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International conference on machine learning, pp. 4904–4916. PMLR, 2021
2021
-
[62]
Zhangqi Jiang, Zheng Sun, Xianfang Zeng, Yufeng Yang, Xuanyang Zhang, Yongliang Wu, Wei Cheng, Gang Yu, Xu Yang, and Bihan Wen. Geditbench v2: A human-aligned benchmark for general image editing.arXiv preprint arXiv:2603.28547, 2026
arXiv 2026
-
[63]
Muon: An optimizer for hidden layers in neural networks, 2024
Keller Jordan, Yuchen Jin, Vlado Boza, Jiacheng You, Franz Cesista, Laker Newhouse, and Jeremy Bernstein. Muon: An optimizer for hidden layers in neural networks, 2024. URLhttps://kellerjordan.github.io/posts/ muon/
2024
-
[64]
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020
Pith/arXiv arXiv 2001
-
[65]
Elucidating the design space of diffusion-based generative models
Tero Karras, Miika Aittala, Samuli Laine, and Timo Aila. Elucidating the design space of diffusion-based generative models. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22, Red Hook, NY, USA, 2022. Curran Associates Inc. ISBN 9781713871088. 49
2022
-
[66]
Analyzing and improving the training dynamics of diffusion models, 2024
Tero Karras, Miika Aittala, Jaakko Lehtinen, Janne Hellsten, Timo Aila, and Samuli Laine. Analyzing and improving the training dynamics of diffusion models, 2024. URLhttps://arxiv.org/abs/2312.02696
Pith/arXiv arXiv 2024
-
[67]
Auto-encoding variational bayes.arXiv preprint arXiv:1312.6114, 2013
Diederik P Kingma and Max Welling. Auto-encoding variational bayes.arXiv preprint arXiv:1312.6114, 2013
Pith/arXiv arXiv 2013
-
[68]
Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advancesin neural information processing systems, 36:36652–36663, 2023
Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advancesin neural information processing systems, 36:36652–36663, 2023
2023
-
[69]
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012
2012
-
[70]
Kling-image-2.1.https://kling.ai/explore/kling_2.1_api, 2025
Kuaishou Technology. Kling-image-2.1.https://kling.ai/explore/kling_2.1_api, 2025
2025
-
[71]
Improved precision and recall metric for assessing generative models.Advances in neural information processing systems, 32, 2019
Tuomas Kynkäänniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, and Timo Aila. Improved precision and recall metric for assessing generative models.Advances in neural information processing systems, 32, 2019
2019
-
[72]
Krea 2.https://www.krea.ai/blog/krea-2-technical-report, 2026
Sangwu Lee, Erwann Millon, Le Zhuo, Matthew Newton, Andrei Filatov, Naga Sai Abhinay Devarinti, Dazhi Zhong, Avram Djordjevic, Gabriel Menezes, Will Beddow, Titus Ebbecke, Mihai Petrescu, Owen Fahey, Gian Saß, Felix Gil, and Victor Perez. Krea 2.https://www.krea.ai/blog/krea-2-technical-report, 2026
2026
-
[73]
Qwen-image-bench: From generation to creation in text-to-image evaluation,
Niantong Li, Guangzheng Hu, Weixu Qiao, Ying Ba, Qichen Hong, Shijun Shen, Jinlin Wang, Fan Zhou, Jianye Kang, Xin Shang, Ziyi He, Wei Wang, Dalin Li, Jiahao Li, Jie Zhang, Kaiyuan Gao, Kun Yan, Lihan Jiang, Ningyuan Tang, Shengming Yin, Tianhe Wu, Xiao Xu, Xiaoyue Chen, Yuxiang Chen, Yan Shu, Yanran Zhang, Yilei Chen, Yixian Xu, Zekai Zhang, Zhendong Wan...
-
[74]
Omnicorpus: A unified multimodal corpus of 10 billion-level images interleaved with text
Qingyun Li, Zhe Chen, Weiyun Wang, Wenhai Wang, Shenglong Ye, Zhenjiang Jin, Guanzhou Chen, Yinan He, Zhangwei Gao, Erfei Cui, Jiashuo Yu, Hao Tian, Jiasheng Zhou, Chao Xu, Bin Wang, Xingjian Wei, Wei Li, Wenjian Zhang, Bo Zhang, Pinlong Cai, Licheng Wen, Xiangchao Yan, Pei Chu, Yi Wang, Min Dou, Changyao Tian, Xizhou Zhu, Lewei Lu, Yushi Chen, Junjun He,...
2025
-
[75]
Reflect-dit: Inference-time scaling for text-to-image diffusion transformers via in-context reflection
Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Arsh Koneru, Yusuke Kato, Kazuki Kozuka, and Aditya Grover. Reflect-dit: Inference-time scaling for text-to-image diffusion transformers via in-context reflection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 15657–15668, 2025
2025
-
[76]
Holistic evaluation of language models.arXiv preprint arXiv:2211.09110, 2022
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. Holistic evaluation of language models.arXiv preprint arXiv:2211.09110, 2022
Pith/arXiv arXiv 2022
-
[77]
Scaling laws for diffusion transformers.arXiv preprint arXiv:2410.08184, 2024
Zhengyang Liang, Hao He, Ceyuan Yang, and Bo Dai. Scaling laws for diffusion transformers.arXiv preprint arXiv:2410.08184, 2024
arXiv 2024
-
[78]
Jintao Lin, Bowen Dong, Weikang Shi, Chenyang Lei, Suiyun Zhang, Rui Liu, and Xihui Liu. Aegis: Exploring the limit of world knowledge capabilities for unified mulitmodal models.arXiv preprint arXiv:2601.00561, 2026
arXiv 2026
-
[79]
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling, 2023. URLhttps://arxiv.org/abs/2210.02747
Pith/arXiv arXiv 2023
-
[80]
Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
2023
-
[81]
Ernie-image technical report.arXiv preprint arXiv:2605.25347, 2026
Jiaxiang Liu, Zhida Feng, Pengyu Zou, Zhenyu Qian, Tianrui Zhu, Jun Xia, Yuehu Dong, Yanzheng Lin, Honglin Xiong, Anqi Chen, et al. Ernie-image technical report.arXiv preprint arXiv:2605.25347, 2026
Pith/arXiv arXiv 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.