REVIEW 2 major objections 6 minor 48 references
Six million Pixiv images show open-source image creators stick to a tiny head of models, lag weeks behind new versions, and get more engagement when they stack LoRAs.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 10:58 UTC pith:72SOJJUV
load-bearing objection First large observational map of real multi-model creative workflows, backed by a public 6 M-image metadata corpus; solid descriptive HCI work with ordinary selection caveats. the 2 major comments →
Navigating the Open-Source Model Ecosystem: An Empirical Study of Creator Practices in Artistic Image Generation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Across 5.98 million Pixiv images the usage of 22.4 thousand base models and 154 thousand LoRAs follows a classic long-tail: the top 2.5 percent of base models generate 80 percent of images, most creators use fewer than five base models, base models peak in roughly eleven weeks while LoRAs peak later and then plateau, only about 40 percent of images adopt the newest version even after twenty weeks, and by 2025 three-quarters of images use at least one LoRA—images that receive systematically more views and bookmarks.
What carries the argument
The generation-metadata recipe embedded in each image file: the exact base-model hash, LoRA hashes, and prompts that let the authors reconstruct the full creative configuration and join it to Civitai model metadata and Pixiv engagement metrics.
Load-bearing premise
That the subset of Pixiv AI artworks whose headers still contain complete, parseable generation metadata (and whose hashes match Civitai) is representative enough of ordinary open-source creator practice to support ecosystem-wide claims.
What would settle it
If an independent crawl of the same period that recovers complete metadata for a substantially larger or differently sampled set of Pixiv AI images yields a markedly flatter usage distribution, faster version adoption, or no engagement premium for multi-LoRA images, the central empirical picture collapses.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents the first large-scale empirical study of how creators select, combine, and update open-source image-generation models (base checkpoints and LoRAs) in real artistic practice. The authors assemble a dataset of ~6 M Pixiv AI-tagged images that retain complete generation metadata, link 22.4 K base models and 154 K LoRAs to Civitai records, and answer three RQs: (1) model-usage diversity and long-tail concentration, (2) model life-cycle and version-update inertia, and (3) LoRA adoption, category shifts, multi-LoRA composition, and correlation with engagement. Key descriptive findings include extreme concentration (top 2.5 % of base models produce 80 % of images), multi-week adoption lags, persistent older-version use, a rise of multi-LoRA workflows that correlate with higher views/bookmarks, and a shift from character LoRAs toward style/concept LoRAs as base models improve. The dataset is released publicly.
Significance. If the observational patterns hold, the work supplies the first systematic, large-scale map of authentic creator practice inside the open-source generative-image ecosystem, a setting that has been studied mainly through model repositories or single-model prompt logs. The public release of millions of image–metadata pairs is a concrete resource for model recommendation, prompt tooling, and socio-technical studies. The documented long-tail concentration, version inertia, and LoRA-composition regularities give platform designers and model authors concrete targets for recommendation systems and compatibility tooling. Strengths that raise the paper’s value are the transparent data pipeline (AI-tag filter, hash matching, public release) and the consistent use of rank plots, entropy, CDFs, and co-occurrence ratios without hidden fitted parameters.
major comments (2)
- §3.1–3.2 and the subsequent RQs rest on the subset of Pixiv AI-tagged artworks whose file headers still contain complete, parseable generation metadata and whose model hashes match Civitai. The paper is transparent about the filter and about the non-negligible “not-matched” fraction (weekly ~47 % base, ~25 % LoRA), yet the central claims about ecosystem-wide diversity, life cycles, and LoRA effectiveness are framed as characterizing the open-source creator community. A short quantitative sensitivity check—e.g., comparing engagement or creator productivity of matched vs. unmatched images, or reporting how many creators contribute only unmatched models—would strengthen the claim that selection does not overturn the headline patterns. Without it the representativeness assumption remains the weakest load-bearing point.
- §6.1 (Figure 11 and the monthly CDFs in the appendix) reports a clear positive association between LoRA count and views/bookmarks, but the analysis is purely correlational. The text sometimes edges toward causal language (“the customization … are correlated to an artwork’s success,” “key benefit”). Because creator skill, prompt quality, and base-model choice are confounded, the manuscript should either (a) add a simple matched comparison (same creator, same month, with/without LoRA) or (b) explicitly restate that the result is associational only and cannot support claims of LoRA “effectiveness.” This is load-bearing for the third contribution listed in the abstract.
minor comments (6)
- Abstract and §1: “wee present” is a typo; correct to “we present.”
- Figure 2 caption and surrounding text: the y-axis label “Number of models” is ambiguous (unique models used that week vs. cumulative); a short clarification would help.
- §4.2: the phrase “normalized Shannon entropy” is used without stating the exact normalization (log of number of active models that week?). A one-line definition would improve reproducibility.
- §5.1: the minimum-usage threshold of 1 000 images for life-cycle curves is reasonable but arbitrary; a brief robustness note (e.g., results for 500 or 2 000) would be useful.
- Figure 14 co-occurrence matrix: the color scale is not shown; adding a legend would make the ratios easier to read.
- Appendix A.1: the 11.2 % preference-for-older-version figure is informative; stating the exact coding protocol (or inter-rater check) would increase confidence.
Circularity Check
No significant circularity: purely observational descriptive statistics from a new Pixiv+Civitai-linked dataset; no fitted parameters re-presented as predictions and no load-bearing self-citation loops.
full rationale
The paper is an empirical measurement study. All headline quantities (long-tail rank-cumulative curves, weekly Shannon entropy, normalized weekly popularity trajectories, peak-reaching CDFs, version-adoption shares, LoRA-count vs. views/bookmarks CDFs, category proportions, and co-occurrence ratios R2/R1) are defined directly from raw image counts, model hashes, and timestamps extracted from the newly constructed 6 M-image corpus. None of these quantities is later re-used as an independent “prediction” of a quantity that was already used to construct it. Self-citations (prior Pixiv AIGC and Civitai descriptive papers by overlapping authors) appear only in Related Work and platform-context paragraphs; they supply background, not uniqueness theorems, ansatzes, or uniqueness claims that force the present results. The selection filter (AI-tagged Pixiv works retaining parseable generation metadata and matchable hashes) is stated transparently in §3 and does not create a definitional loop. Consequently the derivation chain is self-contained observational description with score 0.
Axiom & Free-Parameter Ledger
free parameters (2)
- minimum usage threshold for life-cycle analysis =
1000 uses
- co-occurrence ratio definition R2/R1
axioms (4)
- domain assumption Embedded generation metadata (model hashes, prompts, parameters) in Pixiv image files accurately records the models actually used by creators.
- domain assumption Civitai model hashes and category labels correctly identify the same models and functional types used on Pixiv.
- domain assumption Pixiv view and bookmark counts are valid proxies for artwork success when comparing LoRA usage groups within the same upload month.
- domain assumption The subset of AI-tagged artworks that retain complete parseable metadata is representative of broader open-source creator practice.
read the original abstract
The open-sourcing of powerful image generation models has created a vibrant ecosystem where creators curate and combine a vast array of community-contributed models. This practice stands in sharp contrast to using closed-source tools like Midjourney. Yet, little is known about these emerging creative workflows. To bridge this gap, this paper presents the first large-scale empirical study of creator model usage behavior within this open-source image generation ecosystem. We construct a novel dataset of 6 million images with their embedded generation metadata -- a detailed recipe of the creation process, including the models used and the prompts. By linking the usage of 22.4K base models and 154K LoRA models to the images, our findings underscore the ecosystem's unique strengths and its inherent obstacles. This provides valuable insights for making this ecosystem more sustainable and innovative. Moreover, we make our dataset publicly available, providing creators with practical references for producing better artworks and researchers to facilitate further studies.
Figures
Reference graph
Works this paper leans on
-
[1]
Hanqun Cao, Cheng Tan, Zhangyang Gao, Yilun Xu, Guangyong Chen, Pheng- Ann Heng, and Stan Z. Li. 2024. A Survey on Generative Diffusion Models. IEEE Transactions on Knowledge and Data Engineering36, 7 (2024), 2814–2830. doi:10.1109/TKDE.2024.3361474
-
[2]
Civitai. 2026. Civitai Changelog. https://civitai.com/changelog
2026
-
[3]
Cover and Joy A
Thomas M. Cover and Joy A. Thomas. 1999.Elements of Information Theory(2nd ed.). Wiley-Interscience, New York, NY, USA
1999
-
[4]
Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah
-
[5]
Diffusion Models in Vision: A Survey .IEEE Transactions on Pattern Analysis & Machine Intelligence45, 09 (Sept. 2023), 10850–10869. doi:10.1109/ TPAMI.2023.3261988
arXiv 2023
-
[6]
Maria-Teresa De Rosa Palmini and Eva Cetinic. 2025. Exploring Language Patterns of Prompts in Text-to-Image Generation and Their Impact on Visual Diversity. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’25). Association for Computing Machinery, New York, NY, USA, 2318–2348. doi:10.1145/3715275.3732158
-
[7]
Prafulla Dhariwal and Alexander Quinn Nichol. 2021. Diffusion Models Beat GANs on Image Synthesis. InAdvances in Neural Information Processing Systems, A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan (Eds.). https: //openreview.net/forum?id=AAWuCvzaVt
2021
-
[8]
Smith, Gianluca Guadagni, Jon Chun, Jordi Cabot, Joseph Marvin Imperial, Juan A
Francisco Eiras, Aleksandar Petrov, Bertie Vidgen, Christian Schroeder De Witt, Fabio Pizzati, Katherine Elkins, Supratik Mukhopadhyay, Adel Bibi, Botos Csaba, Fabro Steibel, Fazl Barez, Genevieve A. Smith, Gianluca Guadagni, Jon Chun, Jordi Cabot, Joseph Marvin Imperial, Juan A. Nolazco-Flores, Lori Landay, Matthew Jackson, Paul Röttger, Philip H.S. Torr...
2024
-
[9]
Frank, Matthew Groh, Laura Herman, Neil Leach, Robert Mahari, Alex “Sandy” Pent- land, Olga Russakovsky, Hope Schroeder, and Amy Smith
Ziv Epstein, Aaron Hertzmann, the Investigators of Human Creativ- ity, Memo Akten, Hany Farid, Jessica Fjeld, Morgan R. Frank, Matthew Groh, Laura Herman, Neil Leach, Robert Mahari, Alex “Sandy” Pent- land, Olga Russakovsky, Hope Schroeder, and Amy Smith. 2023. Art and the science of generative AI.Science380, 6650 (2023), 1110–
2023
-
[10]
arXiv:https://www.science.org/doi/pdf/10.1126/science.adh4451 doi:10. 1126/science.adh4451
-
[11]
Ahmad Faisal Choiril Anam Fathoni. 2023. Leveraging Generative AI Solutions in Art and Design Education: Bridging Sustainable Creativity and Fostering Aca- demic Integrity for Innovative Society. InE3S Web Conf., Vol. 426. EDP Sciences, 01102. doi:10.1051/e3sconf/202342601102 The 5th International Conference of Biospheric Harmony Advanced Research (ICOBAR 2023)
-
[12]
Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit Haim Bermano, Gal Chechik, and Daniel Cohen-Or. 2023. An Image is Worth One Word: Per- sonalizing Text-to-Image Generation using Textual Inversion. InThe Eleventh International Conference on Learning Representations. https://openreview.net/ forum?id=NAQvF08TcyG
2023
-
[13]
Shalmoli Ghosh, Matthew R. DeVerna, and Filippo Menczer. 2026. A Marketplace for AI-Generated Adult Content and Deepfakes. arXiv:2601.09117 [cs.CY] https: //arxiv.org/abs/2601.09117
Pith/arXiv arXiv 2026
-
[14]
Yuanhe Guo, Haoming Liu, and Hongyi Wen. 2024. Gemrec: Towards generative model recommendation. InProceedings of the 17th ACM International Conference on Web Search and Data Mining. 1054–1057
2024
-
[15]
Yuanhe Guo, Linxi Xie, Zhuoran Chen, Kangrui Yu, Ryan Po, Guandao Yang, Gordon Wetzstein, and Hongyi Wen. 2025. ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model Personalization. InProceedings of the IEEE/CVF International Conference on Computer Vision. 19577–19586
2025
-
[16]
Will Hawkins, Brent Mittelstadt, and Chris Russell. 2025. Deepfakes on Demand: The rise of accessible non-consensual deepfake image generators. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’25). Association for Computing Machinery, New York, NY, USA, 1602–1614. doi:10.1145/3715275.3732107
-
[17]
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Liang Wang, Weizhu Chen, et al. 2022. Lora: Low-rank adaptation of large language models.Iclr1, 2 (2022), 3
2022
-
[18]
Chengyou Jia, Changliang Xia, Zhuohang Dang, Weijia Wu, Hangwei Qian, and Minnan Luo. 2025. Chatgen: Automatic text-to-image generation from freestyle chatting. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13284–13293
2025
-
[19]
Dongfu Jiang, Max Ku, Tianle Li, Yuansheng Ni, Shizhuo Sun, Rongqi Fan, and Wenhu Chen. 2024. Genai arena: An open evaluation platform for generative models.Advances in Neural Information Processing Systems37 (2024), 79889– 79908
2024
-
[20]
Alireza Joonbakhsh, Alireza Rostami, AmirMohammad Kamalinia, Ali Nazeri, Far- shad Khunjush, Bedir Tekinerdogan, and Siamak Farshidi. 2025. Evidence-Driven Decision Support for AI Model Selection in Research Software Engineering.arXiv preprint arXiv:2512.11984(2025)
arXiv 2025
-
[21]
2024.Does Generative AI Crowd Out Human Creators? Evidence from Pixiv
Sueyoul Kim, Ginger Zhe Jin, and Eungik Lee. 2024.Does Generative AI Crowd Out Human Creators? Evidence from Pixiv. Working Paper 32711. National Bureau of Economic Research. doi:10.3386/w32711
-
[22]
Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. 2023. Pick-a-pic: An open dataset of user preferences for text-to- image generation.Advances in neural information processing systems36 (2023), 36652–36663
2023
-
[23]
Sebastian Krakowski. 2025. Human-AI agency in the age of generative AI.Inf. Organ.35, 1 (March 2025), 25 pages. doi:10.1016/j.infoandorg.2025.100560
-
[24]
Ohchan Kwon and Yue Zhang. 2024. How Do Human Creators Compete with Art-ificial Intelligence? An Empirical Analysis. InPACIS 2024 Proceedings. https: //aisel.aisnet.org/pacis2024/track01_aibussoc/track01_aibussoc/16 16
2024
-
[25]
Please, don’t kill the only model that still feels human
Huiqian Lai. 2026. "Please, don’t kill the only model that still feels human": Understanding the# Keep4o Backlash.arXiv preprint arXiv:2602.00773(2026)
arXiv 2026
-
[26]
Vivian Liu and Lydia B Chilton. 2022. Design Guidelines for Prompt Engineering Text-to-Image Generative Models. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems(New Orleans, LA, USA)(CHI ’22). Association for Computing Machinery, New York, NY, USA, Article 384, 23 pages. doi:10.1145/3491102.3501825
-
[27]
Midjourney. 2026. Midjourney: Official Documentation. https://docs.midjourney. com/hc/en-us
2026
-
[28]
Philipp Moeßner and Heike Adel. 2025. Human VS. AI: A novel benchmark and a comparative study on the detection of generated images and the impact of prompts. InProceedings of the 1stWorkshop on GenAI Content Detection (GenAIDe- tect). 47–58
2025
-
[29]
Jonas Oppenlaender. 2024. A taxonomy of prompt modifiers for text-to-image generation.Behaviour & Information Technology43, 15 (2024), 3763–3776. arXiv:https://doi.org/10.1080/0144929X.2023.2286532 doi:10.1080/0144929X.2023. 2286532
-
[30]
Maria-Teresa De Rosa Palmini, Laura Wagner, and Eva Cetinic. 2025. Civiverse: A Dataset for Analyzing User Engagement with Open-Source Text-to-Image Yiluo Wei, Yupeng He, Qiming Ye, and Gareth Tyson Models. InComputer Vision – ECCV 2024 Workshops, Alessio Del Bue, Cristian Canton, Jordi Pont-Tuset, and Tatiana Tommasi (Eds.). Springer Nature Switzer- land...
2025
-
[31]
Pixiv. 2022. Pixiv Policy for AIGC. www.pixiv.net/info.php?id=8728
2022
-
[32]
Latha Ramamoorthy. 2025. Evaluating generative ai: Challenges, methods, and future directions.International Journal of Future Multidisciplinary Research7, 1 (2025)
2025
-
[33]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-Resolution Image Synthesis With Latent Diffusion Mod- els. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10684–10695
2022
-
[34]
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. 2023. DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation . In2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE Computer Society, Los Alamitos, CA, USA, 22500–22510. doi:10.1109/CVPR52729.2023.02155
-
[35]
2022.lora: Using Low-rank adaptation to quickly fine-tune diffusion models
Simo Ryu. 2022.lora: Using Low-rank adaptation to quickly fine-tune diffusion models. https://github.com/cloneofsimo/lora
2022
-
[36]
Shahar Sarfaty, Adi Haviv, Uri Hacohen, Niva Elkin-Koren, Roi Livni, and Amit H Bermano. 2025. CARLoS: Retrieval via Concise Assessment Representation of LoRAs at Scale.arXiv preprint arXiv:2512.08826(2025)
arXiv 2025
-
[37]
semrush. 2026. Pixiv Statistics. www.semrush.com/website/pixiv.net/overview/
2026
-
[38]
Yang Song and Stefano Ermon. 2020. Generative Modeling by Estimating Gradi- ents of the Data Distribution. arXiv:1907.05600 [cs.LG]
Pith/arXiv arXiv 2020
-
[39]
Keqiang Sun, Junting Pan, Yuying Ge, Hao Li, Haodong Duan, Xiaoshi Wu, Renrui Zhang, Aojun Zhou, Zipeng Qin, Yi Wang, et al . 2023. Journeydb: A benchmark for generative image understanding.Advances in neural information processing systems36 (2023), 49659–49678
2023
-
[40]
WAI-illustrious-SDXL. 2025. Characters Tested and Confirmed to be Sup- ported by WAI-illustrious-SDXL. huggingface.co/datasets/sieecc/WAI-NSFW- illustrious-SDXL/blob/main/4400%2Bcharacters(Already%20tested).txt
2025
-
[41]
Zijie J Wang, Evan Montoya, David Munechika, Haoyang Yang, Benjamin Hoover, and Duen Horng Chau. 2022. Diffusiondb: A large-scale prompt gallery dataset for text-to-image generative models.arXiv preprint arXiv:2210.14896(2022)
Pith/arXiv arXiv 2022
-
[42]
Yiluo Wei and Gareth Tyson. 2024. Understanding the Impact of AI-Generated Content on Social Media: The Pixiv Case. InProceedings of the 32nd ACM Interna- tional Conference on Multimedia(Melbourne VIC, Australia)(MM ’24). Association for Computing Machinery, New York, NY, USA, 6813–6822. doi:10.1145/3664647. 3680631
doi:10.1145/3664647 2024
-
[43]
Yiluo Wei, Yiming Zhu, Pan Hui, and Gareth Tyson. 2024. Exploring the Use of Abusive Generative AI Models on Civitai. InProceedings of the 32nd ACM International Conference on Multimedia(Melbourne VIC, Australia)(MM ’24). Association for Computing Machinery, New York, NY, USA, 6949–6958. doi:10. 1145/3664647.3681052
arXiv 2024
-
[44]
Ruijie Xu, Zengzhi Wang, Run-Ze Fan, and Pengfei Liu. 2024. Benchmarking Benchmark Leakage in Large Language Models. arXiv:2404.18824 [cs.CL] https: //arxiv.org/abs/2404.18824
Pith/arXiv arXiv 2024
-
[45]
Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. 2023. Diffusion Models: A Comprehensive Survey of Methods and Applications.ACM Comput. Surv.56, 4, Article 105 (Nov. 2023), 39 pages. doi:10.1145/3626235
doi:10.1145/3626235 2023
-
[46]
Chenshuang Zhang, Chaoning Zhang, Mengchun Zhang, In So Kweon, and Junmo Kim. 2024. Text-to-image Diffusion Models in Generative AI: A Survey. arXiv:2303.07909 [cs.CV] https://arxiv.org/abs/2303.07909
Pith/arXiv arXiv 2024
-
[47]
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Adding Conditional Con- trol to Text-to-Image Diffusion Models. In2023 IEEE/CVF International Conference on Computer Vision (ICCV). 3813–3824. doi:10.1109/ICCV51070.2023.00355
-
[48]
V10 feels like step back from V9. Too cartonish
Lirui Zhao, Yue Yang, Kaipeng Zhang, Wenqi Shao, Yuxin Zhang, Yu Qiao, Ping Luo, and Rongrong Ji. 2024. Diffagent: Fast and accurate text-to-image api selection with large language model. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6390–6399. Navigating the Open-Source Model Ecosystem: An Empirical Study of Creator...
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.